Large scale loss-of-function mutations during chicken evolution and domestication
Curation statements for this article:-
Curated by eLife
eLife Assessment
This useful study reports potential loss-of-function variants, pseudogenes and gene presence-absence variation across multiple chicken genomes, with potential implications for understanding genome evolution and domestication. The evidence for the central claims is unfortunately incomplete, as the inferences of gene loss are not sufficiently robust to account for assembly and annotation artifacts, and, in addition, the analyses can not distinguish between positive selection and relaxed constraint. The overall claim of large-scale gene loss being adaptive and thus being a major driver of chicken evolution and domestication is therefore not sufficiently supported. The area of the study is of interest to colleagues in evolutionary and comparative genomics as well as animal domestication.
This article has been Reviewed by the following groups
Listed in
- Evaluated articles (eLife)
Abstract
Evolution and domestication are often driven by genetic innovations, yet the role of gene loss remains debated. Here, we present comparative genomic analyses of four indigenous chicken breeds from Yunnan Province, China and a reference red jungle fowl genome (GRCg6a). We identify extensive gene presence–absence variation and large numbers of pseudogenes, revealing highly dynamic gene repertoires among closely related chickens. By reconstructing ancestral gene content, we estimate that the common ancestor harbored at least 21,972 genes, of which 7,993 are dispensable. Each lineage has independently lost thousands of genes through both complete gene loss and pseudogenization. These loss-of-function events are non-random: pseudogenization mutations are biased toward gene termini, frequently fixed in populations, and enriched in specific biological pathways. Notably, patterns of gene loss recapitulate phylogenetic relationships, suggesting that loss-of-function mutations are shaped by selection rather than neutral drift. Analysis of four other chicken genomes assembled using PacBio HiFi reads draws the same conclusion. Thus, our results support a model in which large-scale loss-of-function mutations are a major driver of chicken evolution and domestication, consistent with the “less-is-more” hypothesis of adaptive evolution.
Article activity feed
-
eLife Assessment
This useful study reports potential loss-of-function variants, pseudogenes and gene presence-absence variation across multiple chicken genomes, with potential implications for understanding genome evolution and domestication. The evidence for the central claims is unfortunately incomplete, as the inferences of gene loss are not sufficiently robust to account for assembly and annotation artifacts, and, in addition, the analyses can not distinguish between positive selection and relaxed constraint. The overall claim of large-scale gene loss being adaptive and thus being a major driver of chicken evolution and domestication is therefore not sufficiently supported. The area of the study is of interest to colleagues in evolutionary and comparative genomics as well as animal domestication.
-
Reviewer #1 (Public review):
Summary:
The authors have assembled the genome of four local chicken breeds from China and analysed their gene content. They come to the conclusion that thousands of genes present in the current reference genome of chicken have become pseudogenized during chicken evolution and domestication. They argue that their study provides strong support for the importance of the "less-is-more" hypothesis for adaptive evolution.
Strengths:
The paper provides medium-quality genome assemblies for four individuals representing four local populations of chickens and analyses their gene content.
Weaknesses:
They have not excluded the possibility that the high rate of putative pseudogenes reflects the presence of errors in gene models, in particular in GC-rich microchromosomes that are challenging to assemble correctly. The …
Reviewer #1 (Public review):
Summary:
The authors have assembled the genome of four local chicken breeds from China and analysed their gene content. They come to the conclusion that thousands of genes present in the current reference genome of chicken have become pseudogenized during chicken evolution and domestication. They argue that their study provides strong support for the importance of the "less-is-more" hypothesis for adaptive evolution.
Strengths:
The paper provides medium-quality genome assemblies for four individuals representing four local populations of chickens and analyses their gene content.
Weaknesses:
They have not excluded the possibility that the high rate of putative pseudogenes reflects the presence of errors in gene models, in particular in GC-rich microchromosomes that are challenging to assemble correctly. The paper contains no genotype-phenotype analysis, which means that the adaptive significance of a high rate of pseudogenization, if it exists, is unknown.
-
Reviewer #2 (Public review):
Summary:
The authors set out to investigate the evolutionary role of gene presence-absence variation and pseudogenization in chicken evolution and domestication. By comparing draft genome assemblies of four indigenous Chinese chicken breeds against the red junglefowl reference genome (GRCg6a) and four PacBio HiFi assemblies, the study proposes that the common ancestor possessed nearly 22,000 genes, and that each domestic lineage independently lost thousands of genes (identifying ~8,000 dispensable genes). The authors conclude that massive loss of function and pseudogenization represent major drivers of chicken evolution under the "less-is-more" hypothesis.
While the concept that gene loss can drive phenotypic diversification during domestication is compelling, the results do not convincingly support the …
Reviewer #2 (Public review):
Summary:
The authors set out to investigate the evolutionary role of gene presence-absence variation and pseudogenization in chicken evolution and domestication. By comparing draft genome assemblies of four indigenous Chinese chicken breeds against the red junglefowl reference genome (GRCg6a) and four PacBio HiFi assemblies, the study proposes that the common ancestor possessed nearly 22,000 genes, and that each domestic lineage independently lost thousands of genes (identifying ~8,000 dispensable genes). The authors conclude that massive loss of function and pseudogenization represent major drivers of chicken evolution under the "less-is-more" hypothesis.
While the concept that gene loss can drive phenotypic diversification during domestication is compelling, the results do not convincingly support the central conclusions. The scale of reported gene loss and the specific patterns of pseudogenization appear to be potentially driven by well-known genome assembly gaps, annotation errors, and sequencing dropouts rather than genuine evolutionary events, and the authors do not provide enough convincing evidence that this is not the case.
Strengths:
(1) The study addresses an important and timely evolutionary question regarding the role of gene loss and loss-of-function variation in animal domestication.
(2) The inclusion of multiple indigenous Chinese chicken breeds alongside high-accuracy PacBio HiFi assemblies provides a valuable comparative genomic dataset.
(3) The authors attempt to evaluate pseudogene transcription using RNA-seq data and check for transcript isoforms that bypass candidate loss-of-function mutations.
Weaknesses:
(1) Unvalidated gene and pseudogene annotations: Long-read and consensus genome assemblies are known to suffer from residual indel errors that create artificial frameshifts and premature stop codons. The frequency of predicted pseudogenes in this study (~3.5%-4.5%) aligns closely with expected baseline annotation error rates. Although the authors state in their Methods that mutations were validated using short reads, this validation is never quantitatively demonstrated or shown in the results. The fact that most proposed pseudogenes are actively transcribed and lack a paralog strongly suggests that many are intact, functional genes affected by sequencing or annotation artifacts.
(2) Assembly gaps and GC-bias mistaken for gene loss: The claim that ancestral chickens possessed ~22,000 protein-coding genes and lost thousands of genes in only ~10,000-50,000 years is inconsistent with the evolutionary conservation of avian genomes. The missing genes are enriched for high GC content and preferentially located on microchromosomes. Avian microchromosomes and dot chromosomes are notoriously GC-rich, repeat-dense, and prone to severe assembly gaps in non-telomere-to-telomere assemblies. The reported gene absences reflect assembly fragmentation and coverage dropouts rather than evolutionary deletions.
(3) Lack of synteny validation: Genuine gene absence requires demonstrating conserved collinear synteny of flanking orthologous genes with an unambiguous sequence deletion at the locus. Relying on sequence alignment or short-read mapping failures across fragmented scaffolds substantially inflates false-positive gene loss calls.
(4) Positional bias of pseudogenization mutations: The observed concentration of pseudogenization mutations in the terminal 10% of coding sequences (the "bathtub" distribution) is characteristic of alignment boundary artifacts and non-canonical translation start/stop annotations, rather than positive selection to disrupt gene ends. Mutations in the terminal 3' region often produce functional proteins with slightly altered C-termini rather than complete loss of function.
(5) Inconsistent terminology: The manuscript alternates between identifying pseudogenes as unitary (lacking a functional paralog in the same genome) and evaluating sequence identity against "parental genes," creating substantial confusion regarding whether loci are duplicated paralogs or orthologous reference genes.
-
Reviewer #3 (Public review):
Summary:
The authors reanalyze genome assemblies of four indigenous chicken breeds from Yunnan Province together with the red jungle fowl reference (GRCg6a), and search for genes with disrupted protein-coding sequences. They catalog candidate pseudogenes and missing genes to estimate that 7,993 of the ancestral genes are dispensable. They characterize the positional distribution of pseudogenization mutations along coding sequences, their fixation in breeds, their estimated ages, and their pathway enrichments. Most results are replicated in four independent PacBio HiFi-based chicken assemblies. From the biased position of pseudogenization mutations toward CDS ends, their frequent fixation, and the phylogenetic signal in gene-loss patterns, the authors conclude that large-scale loss of function is a major …
Reviewer #3 (Public review):
Summary:
The authors reanalyze genome assemblies of four indigenous chicken breeds from Yunnan Province together with the red jungle fowl reference (GRCg6a), and search for genes with disrupted protein-coding sequences. They catalog candidate pseudogenes and missing genes to estimate that 7,993 of the ancestral genes are dispensable. They characterize the positional distribution of pseudogenization mutations along coding sequences, their fixation in breeds, their estimated ages, and their pathway enrichments. Most results are replicated in four independent PacBio HiFi-based chicken assemblies. From the biased position of pseudogenization mutations toward CDS ends, their frequent fixation, and the phylogenetic signal in gene-loss patterns, the authors conclude that large-scale loss of function is a major driver of chicken evolution and domestication, consistent with the "less-is-more" hypothesis.
Strengths:
The catalog itself is a substantial resource: the comparison spans multiple closely related genomes, the main patterns are checked in a second, independently assembled set of HiFi genomes, and population resequencing data are used to ask whether the pseudogenization mutations are fixed rather than segregating. The finding that candidate loss-of-function genes are under relaxed purifying selection is well supported. The question of what gene loss contributes to domestication is worth asking, and this is a useful dataset for asking it.
Weaknesses:
The finding that genes carrying disruptive mutations are under relaxed selection is not particularly surprising, and the more interesting claim, that the observed patterns reflect positive selection for gene loss, is less certain in my opinion.
(1) Relaxed purifying selection versus positive selection. The bias of pseudogenization mutations toward the two ends of coding sequences is interpreted as positive or artificial selection, along with elevated dN/dS. But the alternative, that disruptive mutations at gene ends are simply better tolerated, is equally consistent with the results. Several mechanisms would produce this pattern under relaxed constraint without positive selection per se: alternative downstream start sites that rescue 5-prime disruptions; the small fraction of protein truncated by 3-prime disruptions; and the enrichment of disordered regions at protein termini.
(2) The functional status of the pseudogenes is assumed, not demonstrated. Genome-scale work cannot be expected to validate individual genes, but the language of the paper should reflect the candidate status of these calls. Nearly all predicted pseudogenes (~95%) were reported as transcribed in multiple tissues. It is possible for a nonfunctional coding sequence to retain intact regulatory sequences, but the observation deserves more attention in the paper, particularly because transcripts carrying premature termination codons can be the targets of nonsense-mediated decay, which is not discussed. These remain candidate pseudogenes defined by the presence of a putatively large-effect mutation (e.g. premature stop or frameshift).
(3) What "missing" means. For a gene to be scored as completely absent, it could be genuinely deleted, or its allele could be diverged enough that annotation and orthology/mapping no longer detect it. These are different phenomena. Related, since most pseudogenization mutations are reported as fixed or nearly fixed in their populations, the history of alleles matters: it is not clear how the authors established which state is derived, and whether the reference sequence assumed to be functional is in fact the "functional" version.
(4) Limited biological insight into domestication. The main biological interpretation rests on hierarchical clustering of dispensable genes followed by ontology enrichment within clusters (Figure 7a), with narrative connections to breed phenotypes. Only 19.6% of dispensable genes have Gene Ontology assignments, and the phenotype links are speculative. The section on the subspecies origin of the GRCg6a reference is only loosely connected to the loss-of-function story, and the population-genetic analysis supporting it is thin as described.
-