GENCODE v45-v50 human gene-count comparison ProteinIQ, 19 September 2026 Question How much did the reference-chromosome protein-coding catalog change between GENCODE v45 and v50, beyond the difference in its headline totals? Sources and scope Six comprehensive CHR GTF annotations, GENCODE v45-v50, GRCh38.p14. Exact download URLs, download timestamps, compressed-file SHA-256 checksums, and the corresponding published statistics are in sources.json. Release dates are from https://www.gencodegenes.org/human/releases.html. Download and compare each release's gencode.vNN.annotation.gtf.gz, not its basic, primary-assembly, or chr_patch_hapl_scaff annotation. Only gene feature rows are counted. Transcripts, exons, and CDS rows do not add genes. Reference-chromosome files are used without adding other sequences. Definition The counted coding set contains gene_type=protein_coding except rows tagged readthrough_gene. Immune-receptor segment biotypes remain outside this set. The script reproduces all six published total-gene, headline protein-coding, and excluded-readthrough counts before producing the comparison outputs. Matching Match gene IDs after removing only the numeric version suffix. Preserve any _PAR_Y suffix; none occurred among gene rows in these six source files. Do not use gene names as matching keys. Stable-ID matching is an operational comparison of annotations, not a reconstruction of biological gene identity. An identifier coding at both endpoints is retained even if its sequence, coordinates, transcripts, name, or version changed in the intervening period. Transition categories For each adjacent release pair and separately for v45 versus v50: - Retained: identifier belongs to the counted coding set at both endpoints. - Entering/leaving, biotype change: ID is present in both annotations but its gene_type changes into/out of protein_coding. - Entering/leaving, readthrough status: ID remains protein_coding but its readthrough_gene flag changes, moving it into/out of the counted set. - Entering/leaving, identifier absent: ID is missing from the other annotation. These categories are mutually exclusive for each directional comparison. Net change = entering minus leaving. Retained plus entering equals the new count; retained plus leaving equals the old count. The script checks both. Endpoint counts must not be obtained by summing adjacent-release events: an identifier may leave and return during the interval. Ambiguity review For each missing identifier, list every same-strand gene span in the other release that intersects its inclusive genomic interval on the same chromosome. This comparison includes introns and all gene biotypes. Overlap is only a flag for review: it does not prove identity, a split, a merger, or common function. In the v45-v50 comparison, 24 of 29 entering missing IDs and all 6 leaving missing IDs overlap another annotation in the opposite endpoint release. The 5 without such overlap are not established discoveries either. Two illustrative records were inspected across all six releases: - PAXX: ENSG00000148362 in v47 and ENSG00000310560 in v48 have the same name, strand, and chr9 span (136992422-136993984). Endpoint ID matching records one exit and one entry despite this continuity of the annotation. - CAST: ENSG00000153113 changes from protein_coding in v46 to lncRNA in v47, while ENSG00000310517 appears as protein_coding with the same name and an overlapping span. This is not evidence that CAST stopped coding. No attempt was made to adjudicate every split, merger, or identifier replacement. Outputs release-counts.csv: six release totals, raw coding biotype counts, excluded readthrough counts, and lncRNA and TEC counts. coding-transitions.csv: adjacent-release and endpoint summaries. changed-coding-identifiers.csv: each changed identifier, before/after states, versioned IDs, names, coordinates, and missing-ID overlap flags. Filter from_release=45 and to_release=50 for the article's endpoint comparison. provenance.json: source manifest, Python version, and analysis-script checksum. Reproduction Use Python 3.11 or newer; no third-party libraries or paid tools are required. Save analyze.py and sources.json together. Download the six files named in sources.json into INPUT_DIR. From the script directory, run: python analyze.py --input-dir INPUT_DIR --output-dir OUTPUT_DIR The script checks input hashes and published totals, then writes the outputs. Interpretation These are calculations from public annotations, not new sequencing data or experimental validation. Changes in catalog membership cannot be equated with genes gained or lost by humans, or with confirmed new biological discoveries. The chart compares categories of observed identifier turnover using horizontal bars so their counts share a zero baseline and long labels remain readable.