prepR4pcm: An R Package for Preparing Data and Trees for Phylo- genetic Comparative Methods

This article has been Reviewed by the following groups

Read the full article

Listed in

Log in to save this article

Abstract

1. Phylogenetic comparative methods require species names in a trait dataset to match tip labels in a phylogenetic tree. Yet this apparently simple prerequisite is often one of the most fragile steps in a comparative workflow. Names may differ because of, for example, formatting, taxonomic revisions, synonyms, or spelling errors. If these differences are resolved informally, species can be lost from analyses, and the reasons for their loss can be difficult to reconstruct. 2. Here, we present prepR4pcm, an R package for preparing data and trees for phylogenetic comparative methods. The package reconciles species names through a staged procedure: exact matching, normalised matching, synonym lookup with local taxonomic databases, and optional fuzzy matching for likely spelling errors. Each decision is stored in a reconciliation object with the original name, matched name, match type, confidence score, and a short explanation. This object turns name matching from a hidden preprocessing step into an auditable part of the analysis. 3. prepR4pcm also supports the points where comparative workflows need human judgement. Users can inspect unresolved names, accept or reject suggested matches, add manual corrections, apply taxonomy crosswalks (which link names across taxonomic systems), compare reconciliation runs, and generate reports. The package then returns a matched data frame and pruned tree with the same species set, ready for phylogenetic generalised least squares, phylogenetic mixed models, phylogenetic meta-analysis, and related workflows. If users do not yet have a tree, prepR4pcm can retrieve trees from several sources, date trees when suitable information is available, and format tree-source citations. 4. We illustrate the workflow using bundled datasets with realistic name mismatches. prepR4pcm is available at https://github.com/itchyshin/prepR4pcm with documentation and vignettes covering data and tree reconciliation, tree retrieval, multi-tree workflows, and phylogenetic meta-analysis.

Article activity feed

  1. The preprint was assessed very positively by two reviewers and I recommend incorporating their minor feedback into a revision. One reviewer suggests a more realistic data set including taxonomy changes. This would certainly strengthen the manuscript and I leave it to the authors to decide how feasible that would be.

  2. Firstly I would like to appologize to the authors and Recommender for the slight delay in providing comments. I enjoyed reading the manuscript and think it is timely contribution. As the authors acknowledge there are indeed R packages available that allow users to do some of the steps involved in matching a datafile to a phylogeny for comparative analyses, but none allow to undertake all necessary steps, in a single package and for very different model systems. More importantly, the main contribution of this proposed package is the full transparency of the steps involved which I think is not only important but also valuable for reproducibility. I think the manuscript overall does a good job presenting the package and its capabilities, acknowledging previous contributions and making the differences between this proposed package and previous ones clear. My only doubt is whether it might not be valuable for readers if a more realistic example was presented. In my opinion the key problem when working with multi-species data is changes in taxonomy, which results in differences in species names across published datasets, and with published phylogenies. The authors actually point to this problem when referring to the Jetz et al bird phylogeny and the AVONET dataset. I think that said public database, the Jets phylogeny and the more recent Mc Tavish et al phylogeny of birds could make for a nice empirical example of how their proposed package can work in combining these three sources of valuable information for posterior comparative analysis. I personally very recently struggled with combining data from different sources and different phylogenies and ended up having to make code of my own, combined with available packages, and would have loved to have access to prepR4pcm for this task, to a large extent because of the possibility of having a document that sumarises the different steps that were involved and traces decisions that might have been necessary. I think a more down-to-earth example would be highly valuable for readers, but I understand that its maybe a big ask and I don't consider this should impede publication of this work as it is already a clear and concise manuscript and valuable contribution. 

    Some very minor suggestions: 

    Ln 32-33 Not sure if including phylogenetic path analysis could be worth it here. 

    Ln 42, I think the mentioned "other cases" are actually the problems that colleagues are likely to face the most and are the most challenging to address. 

    Ln 48-49 I believe something is missing at the end of the sentence it seems to have gotten truncated. Also I think a citation might be needed. 

    Ln 73 there is a typo "heck" should be check

    Ln 138 I suggest to replace This object with The mapping table...

    In section 2.4 I suggest to make the point more clearly that splitting / lumping creates a problem that automatically requires user judgement, and that needs to be dutifully documented for transparency. 

    Ln 156-157 in addition to adding species close to a relative, another very valid option is simply to replace a closely related species present in the tree for which there is no data, with one for which data is available but is not present in the tree, as this at the end has no effect on the co-variance matrix. 

    I look forward to seeing this work published and are happy the authors chose to do so in PCI!

  3. I have read the manuscript entitled "prepR4pcm: An R Package for Preparing Data and Trees for Phylogenetic Comparative Methods" by Nakagawa et al. This manuscript is a technical note introducing a new R package to handle a precise issue (species labels mismatch) in comparative analysis.

    The case for the usefulness of the package is convincing. However, the negative impact of wrong label mismatch handling is not quantitatively measured, though this would probably be beyond the format of a technical note. The main argument here are that (i) label mismatchs can have (mainly unmeasured as mentioned above) negative impacts, and (ii) it can only help to make this step more robust and, overall, more auditable. A last argument is of course that no R package currently and thoroughly handle the precise use-case here. I agree these arguments are sufficient to justify this package and acceptance of this manuscript with minor revisions detailed below (and provided PCI Evol Biol is OK with accepting such technical notes).

    From a more formal point of view, the manuscript is well-written, well-structured and clear enough. I have only very minor comments along the text :
    - l.107: The rest seems obvious enough, but I think more information about how prepR4pcm "deals with [...] infra-specific rank abbreviation" is needed. What is meant exactly by "abbreviation" here? Are those ignored? Does it mean that sub-species levels are not handled by the package? A bit tangential, but how are mentions like "sp." for defined genus, but undefined species handled?
    - l.112-114: The two sentences are contradictory. We are informed that synonym matches are given lower score, but then that their score is 1.0. Is this a typo?
    - l.135: Note that the comas are in mono font, while they should be in plain type font.
    - l.172, 184, 186: Issue with the formatting of the commands in the paragraph alignment.
    - l.202-203: Please indent the code to improve readability. Also, on the last command, the tilde is not a "short tilde" (U+007E) but a "long tilde operator" (U+223C, or so I think), and so the command does not paste properly in R (which expects the former, not the latter).
    - l.232-233: The command complains first about rotl not being installed, and after installed rotl, the second run complains about fishtree not being installed. Either those packages should be dependencies of the prepR4pcm package, or the code in the manuscript should mention that the need to be installed.


    Title and abstract

        Does the title clearly reflect the content of the article? [X] Yes, [ ] No (please explain), [ ] I don’t know
        Does the abstract present the main findings of the study? [X] Yes, [ ] No (please explain), [ ] I don’t know

    Introduction

        Are the research questions/hypotheses/predictions clearly presented? [X] Yes, [ ] No (please explain), [ ] I don’t know
        Does the introduction build on relevant research in the field? [X] Yes, [ ] No (please explain), [ ] I don’t know

    Materials and methods

        Are the methods and analyses sufficiently detailed to allow replication by other researchers? [X] Yes, [ ] No (please explain), [ ] I don’t know
        If applicable (for empirical studies), are sample sizes are clearly justified? [ ] Yes, [ ] No (please explain), [X] Not applicable
        Are the methods and statistical analyses appropriate and well described? [X] Yes, [ ] No (please explain), [ ] I don’t know

    Results

        In the case of negative results, is there a statistical power analysis (or an adequate Bayesian analysis or equivalence testing)? [ ] Yes, [ ] No (please explain), [X] Not applicable
        Are the results described and interpreted correctly? [ ] Yes, [ ] No (please explain), [ ] Not applicable

    Discussion

        Have the authors appropriately emphasized the strengths and limitations of their study/theory/methods/argument? [X] Yes, [ ] No (please explain), [ ] I don’t know
        Are the conclusions adequately supported by the results (without overstating the implications of the findings)? [X] Yes, [ ] No (please explain), [ ] I don’t know

  4. Phylogenetic comparative methods are widely used in evolutionary biology. In these analyses, reconciling species names between trait datasets and phylogenetic trees is crucial, but an often poorly documented source of error. This study addresses this practical but fundamental step in comparative analyses. The staged approach implemented in prepR4pcm, combining exact and normalised matching, synonym resolution, optional fuzzy matching and explicit user intervention, provides a useful framework for dealing with taxonomic inconsistencies while retaining a record of the decisions made.

    Although there are other R packages available that allow users to do some of the steps involved in matching a datafile to a phylogeny for comparative analyses, none allows the user to undertake all necessary steps in a single package. I particularly appreciated that the package makes data–tree reconciliation an auditable part of the analysis rather than an informal preprocessing step. This approach facilitates transparency and reproducibility and makes it easier to understand why particular species were retained or excluded. The integration of tree retrieval and preparation further increases its practical utility for comparative workflows. The peer-review process emphasized the importance of a realistic and complex data set and the authors decided to keep a simpler data set as an example in the text and point to vignettes with complex examples.

    References

    Shinichi Nakagawa, Santiao Ortega, Ayumi Mizuno, Eduardo Santos, Malgorzata Lagisz, Bhavya Jain, Jimuel Jr Celeste, Sergio Poo Hernandez (2026) prepR4pcm: An R Package for Preparing Data and Trees for Phylo- genetic Comparative Methods. Ecology and Evolutionary Biology, ver.3 peer-reviewed and recommended by PCI Evolutionary Biology https://doi.org/10.32942/X2468Z