GFF3 vs. GTF vs. BED: Genome Annotation Formats Explained Simply
When downloading genome annotation data from Ensembl, NCBI, or the UCSC Genome Browser, researchers are confronted with multiple file formats: GFF3, GTF, and BED. In most tool documentations, the differences between these formats are not explained clearly. They are presented simply as alternative download options, implying they are interchangeable.
This lack of clarity is a significant source of pipeline failures. These three files are not interchangeable: they encode biological coordinates and feature relationships using fundamentally different mathematical and structural rules. Choosing the wrong format for a task, or converting between them using incorrect assumptions, routinely introduces off-by-one coordinate errors, breaks transcript assembly pipelines, or silently erases critical gene-hierarchy metadata.
This article serves as a definitive, technically precise reference guide to genome annotation formats. We define each format structurally, analyze their distinct coordinate systems, explore their hierarchical models, and outline standard conversion workflows.
Comparing Genome Annotation Formats
The different architectures are optimized for different computational tasks, transitioning from highly structured gene hierarchies to flat region intervals:

What Genome Annotation Actually Is
Before examining the differences, it is useful to establish their shared purpose. A genome annotation file contains structured records of where specific biological features (such as genes, transcripts, exons, UTRs, and regulatory elements) are physically located on a reference genome. It records their chromosome coordinates, strand orientation, and functional attributes, providing the structural map required to make sense of raw sequencing reads.
1. GFF3 (General Feature Format Version 3)
GFF3 represents the most expressive and robust standard of the GFF format family, designed to capture complex, multi-level biological hierarchies.
Structural Rules:
- 9 Tab-Delimited Columns: Every GFF3 record is represented by a single line containing nine strict columns:
seqid(Chromosome or scaffold ID).source(The program or database that generated the feature).type(The sequence ontology term, e.g., gene, mRNA, exon).start(1-based integer start coordinate).end(1-based integer end coordinate, inclusive).score(Floating-point value representing confidence or expression).strand(Orientation:+,-,., or?).phase(For CDS features:0,1,2, or.).attributes(A flexible, semicolon-separated list of tag-value pairs, e.g.,ID=gene1;Name=BRCA1).
- The Hierarchical Parent Model: GFF3 defines multi-level biological hierarchies explicitly. It uses
IDandParenttags in the 9th column to link features (e.g., an exon record containsParent=mRNA1, and the mRNA record containsParent=gene1). - The Coordinate System: GFF3 utilizes a 1-based coordinate system with closed intervals, meaning the first nucleotide of a chromosome is coordinate 1, and the interval includes both the start and end coordinates.
- When to Choose It: Choose GFF3 when you need to load full, structured gene models into genome browsers (such as IGV or JBrowse), or when feeding annotations into tools (like MAKER) that require explicit, multi-level parent-child relationships. Its main downside is that attribute formatting can vary slightly between databases (NCBI vs. Ensembl), sometimes requiring custom parsing.
2. GTF (Gene Transfer Format / GFF2)
GTF is a highly popular, more constrained variant of the GFF family, developed specifically to standardize gene annotations for transcriptomic pipelines.
Structural Rules:
- Mandatory Standardized Attributes: While it shares the first 8 columns with GFF3, GTF enforces strict formatting rules in the 9th column. Every record must contain
gene_idandtranscript_idattributes, formatted as space-separated, double-quoted values (e.g.,gene_id "gene1"; transcript_id "tx1";). - Implicit Hierarchy: GTF does not use explicit
Parenttags. Instead, the relationship between features is inferred implicitly: any exon record sharing the identicaltranscript_idis assumed to belong to that transcript, and any transcript sharing the samegene_idbelongs to that gene. - The Coordinate System: Like GFF3, GTF uses a 1-based, closed-interval coordinate system.
- When to Choose It: GTF is the absolute standard for RNA-seq quantification and alignment pipelines. Tools like STAR, HISAT2, HTSeq, featureCounts, and StringTie expect GTF format. Feeding these tools a GFF3 file often causes them to crash or misinterpret the exon-to-transcript relationships.
3. BED (Browser Extensible Data)
BED is structurally distinct from the GFF family, designed not as a highly structured gene model format, but as a general-purpose interval format for region calculations.
Structural Rules:
- Flexible Column Structure: A BED file contains up to 12 columns. Only the first 3 are required:
chrom(Chromosome name).chromStart(The start coordinate).chromEnd(The end coordinate). Optional columns add name, score, strand, thick-start/end, and block counts to represent multi-exon structures.
- The Coordinate System: Crucially, BED utilizes a 0-based coordinate system with half-open intervals. Under this rule, the first nucleotide of a chromosome is coordinate 0, and the end coordinate is non-inclusive (the interval includes the start coordinate but excludes the end coordinate).
- When to Choose It: BED is the format of choice for interval arithmetic and genomic region calculations (such as intersections, subtractions, or coverage overlaps) in utilities like BEDTools or pybedtools. It is also the standard format for peak callers (ChIP-seq and ATAC-seq) and for defining genomic target regions. It is not designed to store complex, nested gene-transcript hierarchies.
The Coordinate Mismatch: 0-Based vs. 1-Based
The difference between 0-based and 1-based coordinate systems is the most common cause of off-by-one errors in computational biology:
Let us illustrate this with a single, concrete example: a biological feature spanning the first 100 base pairs of a chromosome.
- In GFF3 / GTF (1-based, closed interval):
chr1 . feature 1 100 . + . attributesThe interval starts at 1 and ends at 100, containing exactly 100 bases (including both 1 and 100). - In BED (0-based, half-open interval):
chr1 0 100 feature_nameThe interval starts at 0 and ends at 100. It includes coordinate 0 but excludes 100, containing exactly 100 bases (bases 0 through 99).
The physical sequence mapped is identical, but the start coordinates differ by 1. Manually converting between these files without adjusting the start position will shift your genomic intervals, disrupting downstream analysis.
Conversion and Interoperability
Because biological pipelines require different formats at different stages, converting between these formats is a standard task. However, because conversion can be lossy, researchers should utilize established, robust tools rather than writing custom parsing scripts:
- gffread (Cufflinks utility): The gold standard for rapid, reliable conversion between GFF3 and GTF formats (e.g.,
gffread annotation.gff3 -T -o annotation.gtf). - AGAT (Automated Genome Annotation Toolkit): A comprehensive suite of tools designed specifically to clean, validate, and convert complex GFF3 and GTF files while repairing broken parent-child hierarchies.
- gtf2bed (UCSC / BEDOPS utilities): Highly reliable scripts to convert GTF gene models into BED12 intervals, automatically handling the 0-based to 1-based coordinate shift.
Summary: The Decision Framework
To avoid formatting errors, select your format based on your active computational task:
- Use GTF when running RNA-seq alignment, transcript assembly, or quantification pipelines (STAR, featureCounts, StringTie).
- Use GFF3 when loading complete, structured gene annotation datasets into genome browsers (IGV, JBrowse) or running annotation workflows (MAKER).
- Use BED when performing interval arithmetic, calculating genomic region overlaps (BEDTools), or defining target regions.
Whenever possible, download the native, pre-compiled format directly from your primary database source (Ensembl, NCBI, or UCSC) rather than converting it yourself, ensuring absolute coordinate consistency and scientific integrity across your pipelines.
References and Authoritative Specifications
For computational biologists seeking to inspect the low-level formatting specifications discussed, the following resources serve as primary references:
- The GFF3 Format Specification: Maintained by the Sequence Ontology group. github.com/The-Sequence-Ontology/Specifications
- The GTF Format Specification: As documented by Ensembl and the Sanger Institute. ensembl.org/info/website/upload/gff
- The BED Format Specification: Maintained by the UCSC Genome Browser database team. genome.ucsc.edu/FAQ/FAQformat
- The gffread Utility: Pertea, G., & Pertea, M. (2020). "GFFread: a versatile tool for analyzing and manipulating GFF/GTF files." F1000Research, 9, 304. doi:10.12688/f1000research.23292.1
- The BEDTools Suite: Quinlan, A. R., & Hall, I. M. (2010). "BEDTools: a flexible suite of utilities for comparing genomic features." Bioinformatics, 26(6), 841-842. doi:10.1093/bioinformatics/btq033
- The BEDOPS Suite: Neph, S., Kuehn, M. S., Reynolds, A. P., et al. (2012). "BEDOPS: high-performance genomic feature operations." Bioinformatics, 28(14), 1919-1920. doi:10.1093/bioinformatics/bts277
- The AGAT Toolkit: Dainat, J. (2022). "Another Gtf/Gff Analysis Toolkit (AGAT): Resolve interoperability issues and accomplish more with your annotations." Zenodo. doi:10.5281/zenodo.3552717
Explore Purna's Molecular Intelligence Platform
AI-powered workspace for biology teams to accelerate drug discovery from target identification to lead optimization.
Try Purna AI →