Monkey hear, monkey do what? An application of Automated Behavioural Response systems for hypothesis testing in the world’s smallest monkey

Curation statements for this article:
  • Curated by eLife

    eLife logo

    eLife Assessment

    This study applies a relatively novel method, ABR (Automated Behavioural Response system), to an arboreal small mammal, providing some valuable results on the responses of these small primates to both natural predators and human stimuli. However, the study in its current form is incomplete in its framing and methodology and could be improved to appeal more broadly to behavioural biologists.

This article has been Reviewed by the following groups

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Abstract

Observer effects are a frequent problem in animal behaviour studies, particularly when assessing responses to human disturbance. Automated Behavioural Response (ABR) systems, which combine camera traps with automated sound playbacks, offer a solution but have been primarily used on large terrestrial mammals. Here, we demonstrate their use in a small (∼110g) arboreal primate, the eastern pygmy marmoset ( Cebuella niveiventris ). We conducted two playback experiments to test the risk-disturbance and distracted prey hypotheses. The marmosets exhibited strong anti-predator responses to avian predator calls, including increased fleeing and vocalisations. Human speech elicited similar but weaker responses, indicating that pygmy marmosets do not perceive raptors and humans as equivalent threats. Embedding predator calls into anthropogenic noise reduced vocal responses, suggesting that anthropogenic noise interferes with responses to predation cues. Across five weeks, we generated 128 successful experimental trials, demonstrating that ABRs can rapidly produce sample sizes sufficient for hypothesis testing in the field.

Article activity feed

  1. eLife Assessment

    This study applies a relatively novel method, ABR (Automated Behavioural Response system), to an arboreal small mammal, providing some valuable results on the responses of these small primates to both natural predators and human stimuli. However, the study in its current form is incomplete in its framing and methodology and could be improved to appeal more broadly to behavioural biologists.

  2. Reviewer #1 (Public review):

    Summary:

    The authors test specific but related hypotheses regarding anti-predator responses of wild marmoset groups to predator and human playback sounds triggered to play via a motion sensor on camera-trap devices. The differential responses they observe to human noises and natural predator sounds are interesting, but greater inferences are limited due to a lack of clarity in the methods and analyses.

    Strengths:

    The authors create an excellent experimental design using a customised ABR system for testing the behavioural responses of wild, social-living, tiny, arboreal primates: pygmy marmosets. Much of the work is described with great transparency, and figures and tables are helpful in facilitating this.

    Weaknesses:

    The current study requires improvement in three areas, in my opinion, to permit readers to better evaluate the validity and importance of these results.

    (1) Improve the framing of the paper:

    The current title and justification for this study appear to point to a lack of previous studies testing specific hypotheses (line 51/52: "ABRs have not been applied to hypothesis testing". I find this a rather strange argument to make, given that a quick read through of other ABR papers, cited by the authors (e.g., Kasper et al., 2025, Epperly et al., 2021), are testing predictions set by ecological theory in the cascading effects of predator-prey dynamics. To me, even if these are not explicitly stating "X hypothesis" in their paper, they still appear to be studies guided by implicit hypotheses. To say that previous work with ABR did not test hypotheses is presumptuous, in my opinion. The entire paper would be much better appreciated if the authors could reframe the study for its importance to arboreal mammal/ tropical ecology, anthropogenic effects, and so on. Similarly, the authors should avoid use of phrasing such as "this study is the first direct test of .... " (lines 59/60) and should emphasize the true significance of their work, beyond it being the 'first' of something.

    Similarly, on line 87, "demonstrating that the ABR system can be used to generate data for hypothesis testing" should be removed, as firstly sufficient sample size for any study depends on a number of study-specific parameters, and the authors do not actually demonstrate this in my opinion, given that many of their models end up suffering from singular fit. This is due to a lack of sample size, and also because they do not actually do any type of power analysis or something similar to demonstrate that they actually assessed sample size. So again, my suggestion is to reframe the paper to focus on the behavioural ecology and conservation-related impacts rather than this emphasis on methodology.

    (2) Methods:

    The authors generally do a great job providing sufficient detail on the ABR system and how each experiment was designed. Still, there is room for improvement, as I was confused a number of times. I also would recommend that the authors include a limitations section somewhere which considers the drawbacks of their study, in particular the lack of individual identity for behavioural responses of marmosets (especially given that they used focals, it seems), the groups being in close vicinity of one another/potentially related (?), the specific stimuli used, etc.

    Points of confusion for me included what the control was for Experiment 1. Line 93 - 70 videos without playbacks are used (Table 1), but it is not clear how these videos were selected, and it is not a suitable control comparison for assessing the difference in behaviour associated with playbacks (playbacks with control sounds are). At most, these videos will give basal rates of behaviour (like vocalizations, etc.), but then it is not clear why these '70 videos' and how they were chosen to avoid bias. So, for example, in line 105 the authors write that focals were more likely to flee when hearing playback stimuli than in these "control" videos, but this is not convincing. If there was no fleeing after playback of control sounds (cicadas, macaws) - i.e., the true control in this experiment - then this should be the comparison that is emphasized.

    Can the authors also clarify how they considered/assessed the sound playback level (normally done in Decibels) and if they did not normalize the sound level across the playback stimuli, why not, and what potential effect this could have on the results (i.e., something else to consider for the limitations section)?

    Something else not discussed is the rate of exposure to predator and human noise for these wild monkeys. Are these rates within normal range/expectation for these monkeys? Thinking here of the number of videos captured for each group presumably means exposure to a playback unless 'control' videos were videos where no playback sound was emitted (see question above re: control videos). There were a lot more unsuccessful videos than successful ones that the authors could use, so trying to understand the potential impacts of this (see question re: trial/video # below as well).

    (3) Analysis:

    A few things are unclear and need more explanation in the way the authors conducted their analyses, although they do well to detail all steps of their statistical methods, which was great.

    For assessing model fit, it is not clear what exactly was assessed with the 'performance package' line 467, as the authors do not go on to provide us with any results of the performance/fit. Instead, they tell us that the models did not fit well, with no parameter provided (e.g. lines 481-488). I'm familiar with overdispersion as a parameter that is reported for Poisson models (that does not seem to be provided here). Or by looking at changes in model estimates if one datapoint (and/or one group) is removed after another (with replacement, so keeps sample size static). On line 468/469, it says that model fit was assessed via conditional R2; could the authors provide a citation for this practice, and then provide the R2 parameter for the other models that were used/included in the end?

    Also, please standardize how the GLMM results are presented. There should be the estimate, SE, Z or t, then p value. (line 99, 137).

    Given the high rates of exposure to playbacks, I think the authors should test trial # (or video #) for a potential habituation effect, with earlier captures more likely to draw stronger responses than later video captures for each group.

  3. Reviewer #2 (Public review):

    Summary:

    The article describes an interesting methodology to test hypotheses about the impact of anthropogenic noise on a small arboreal primate, the pygmy marmoset. The authors used a motion-triggered combination of camera traps and speakers to play back control sounds, avian predator calls, and anthropogenic noise to test the risk-disturbance hypothesis and the distracted prey hypothesis. In addition, the authors implemented a technique that is usually used for larger mammals and has not been used before for smaller arboreal animals. The authors are careful in their interpretation of the results and do not favor one hypothesis over the other. The authors also elaborate extensively in their discussion on how to improve this kind of data collection in the future.

    Strengths:

    This study provides a method for rapid data collection while minimizing observer impact. The sample size is comparatively large for a wild animal in a reserve, given the overall observation time. The article also benefits from a solid analysis of the data.

    Weaknesses:

    Though the authors tested two contrasting hypotheses, the discussion would benefit from more detail on the ecological relevance of the observed behaviors.

  4. Author response:

    Reviewer #1 (Public review):

    Summary:

    The authors test specific but related hypotheses regarding anti-predator responses of wild marmoset groups to predator and human playback sounds triggered to play via a motion sensor on camera-trap devices. The differential responses they observe to human noises and natural predator sounds are interesting, but greater inferences are limited due to a lack of clarity in the methods and analyses.

    Strengths:

    The authors create an excellent experimental design using a customised ABR system for testing the behavioural responses of wild, social-living, tiny, arboreal primates: pygmy marmosets. Much of the work is described with great transparency, and figures and tables are helpful in facilitating this.

    Weaknesses:

    The current study requires improvement in three areas, in my opinion, to permit readers to better evaluate the validity and importance of these results.

    (1) Improve the framing of the paper:

    The current title and justification for this study appear to point to a lack of previous studies testing specific hypotheses (line 51/52: "ABRs have not been applied to hypothesis testing". I find this a rather strange argument to make, given that a quick read through of other ABR papers, cited by the authors (e.g., Kasper et al., 2025, Epperly et al., 2021), are testing predictions set by ecological theory in the cascading effects of predator-prey dynamics. To me, even if these are not explicitly stating "X hypothesis" in their paper, they still appear to be studies guided by implicit hypotheses. To say that previous work with ABR did not test hypotheses is presumptuous, in my opinion. The entire paper would be much better appreciated if the authors could reframe the study for its importance to arboreal mammal/ tropical ecology, anthropogenic effects, and so on. Similarly, the authors should avoid use of phrasing such as "this study is the first direct test of .... " (lines 59/60) and should emphasize the true significance of their work, beyond it being the 'first' of something.

    Similarly, on line 87, "demonstrating that the ABR system can be used to generate data for hypothesis testing" should be removed, as firstly sufficient sample size for any study depends on a number of study-specific parameters, and the authors do not actually demonstrate this in my opinion, given that many of their models end up suffering from singular fit. This is due to a lack of sample size, and also because they do not actually do any type of power analysis or something similar to demonstrate that they actually assessed sample size. So again, my suggestion is to reframe the paper to focus on the behavioural ecology and conservation-related impacts rather than this emphasis on methodology.

    We agree that the reviewer makes a valid point that while we were focusing on explicit hypothesis testing in our statements that the other papers mentioned are making predictions based off ecological theory. We will remove line 87 and make sure to limit these comments in the manuscript. We will reframe the introduction and discussion to better reflect this and change the verbiage throughout to focus on behavioural ecology, the impacts of anthropogenic noise and the conservation implications of the study as well as the novel arboreal aspect of the work. We will also update the title to better fit this framing of the study.

    (2) Methods:

    The authors generally do a great job providing sufficient detail on the ABR system and how each experiment was designed. Still, there is room for improvement, as I was confused a number of times. I also would recommend that the authors include a limitations section somewhere which considers the drawbacks of their study, in particular the lack of individual identity for behavioural responses of marmosets (especially given that they used focals, it seems), the groups being in close vicinity of one another/potentially related (?), the specific stimuli used, etc.

    We do touch on the limitation of not identifying individuals in the discussion (lines 252-258), but we will draw this into a specific section which will address this and the other limitations mentioned here.

    Points of confusion for me included what the control was for Experiment 1. Line 93 - 70 videos without playbacks are used (Table 1), but it is not clear how these videos were selected, and it is not a suitable control comparison for assessing the difference in behaviour associated with playbacks (playbacks with control sounds are). At most, these videos will give basal rates of behaviour (like vocalizations, etc.), but then it is not clear why these '70 videos' and how they were chosen to avoid bias. So, for example, in line 105 the authors write that focals were more likely to flee when hearing playback stimuli than in these "control" videos, but this is not convincing. If there was no fleeing after playback of control sounds (cicadas, macaws) - i.e., the true control in this experiment - then this should be the comparison that is emphasized.

    For one group, there were only 13 videos without a playback where a marmoset was present. So, we used these 13 videos for this group, and selected 13 videos at random from the other groups to match sample size. For these groups, we assigned each video without a playback but with a pygmy marmoset a sequential number, and then used a random number generator to select 13 videos for analysis. We refer to these videos as controls as they are negative controls, and the reviewer is correct – they do measure basal levels of behaviour. In contrast, the playback of control sounds is a procedural control (Bui et al., 2022). We will change how we refer to these controls in the manuscript to reflect the types of control they are. We use the comparison between negative controls and videos with playbacks in experiment 1 to assess the impact of the playback procedure itself, though the reviewer is correct that our conclusions would be better supported if we explicitly compared the intervention playbacks with the procedural control. We will add post-hoc tests to make this comparison explicit.

    Can the authors also clarify how they considered/assessed the sound playback level (normally done in Decibels) and if they did not normalize the sound level across the playback stimuli, why not, and what potential effect this could have on the results (i.e., something else to consider for the limitations section)?

    We edited the audios so that they were at similar volume levels. We will update the methods with this information and touch on this in the updated limitations section discussed above.

    Something else not discussed is the rate of exposure to predator and human noise for these wild monkeys. Are these rates within normal range/expectation for these monkeys?

    Although we do not have information about exposure rates to predators, we do mention levels of exposure of human noise in our methods section on lines 318-320 and we also touch on this in our discussion lines 206-212.

    We will update the methods section to be clearer that all groups are exposed to high levels of anthropogenic noise due to their proximity to the ecotourism lodge and community. We will also expand on this in the discussion as a limitation as having a broader array of groups with varying levels of exposure to humans would allow us to see the broader behavioural reactions to these playback stimuli.

    For the predators we mention in the methods section line 382 “All four species have been found in the study area (Barker and Papworth, 2024)” but we will further expand on this to provide information on the density of raptors in the area based on our previous study.

    Thinking here of the number of videos captured for each group presumably means exposure to a playback unless 'control' videos were videos where no playback sound was emitted (see question above re: control videos). There were a lot more unsuccessful videos than successful ones that the authors could use, so trying to understand the potential impacts of this (see question re: trial/video # below as well).

    The number of videos was the total number of times the camera trap triggered. These included videos triggered by another animal or foliage movement where there was no marmoset present, and includes both videos with and without a playback.

    We will make this clearer in our description of the results and will report the number of unsuccessful playbacks.

    (3) Analysis:

    A few things are unclear and need more explanation in the way the authors conducted their analyses, although they do well to detail all steps of their statistical methods, which was great.

    For assessing model fit, it is not clear what exactly was assessed with the 'performance package' line 467, as the authors do not go on to provide us with any results of the performance/fit. Instead, they tell us that the models did not fit well, with no parameter provided (e.g. lines 481-488). I'm familiar with overdispersion as a parameter that is reported for Poisson models (that does not seem to be provided here). Or by looking at changes in model estimates if one datapoint (and/or one group) is removed after another (with replacement, so keeps sample size static). On line 468/469, it says that model fit was assessed via conditional R2; could the authors provide a citation for this practice, and then provide the R2 parameter for the other models that were used/included in the end?

    We used various tests from the performance package (e.g. check_overdispersion) to test the fit of different distributional models (e.g. Poisson, negative binomial) to the same data. Although some models were not overdispersed and did not show evidence of zero-inflation, they did have singular fits, and/or a conditional R2 of 1.0 suggesting overfitting. We therefore did not choose these models. We will clarify this and provide further details of our approach in the manuscript.

    The models for the behaviours not reported did not fit well using any distributional model. These behaviours had very low occurrence (0 seconds in most videos), and so there was very sparse data for generating estimates, and very low variation within / between groups and conditions. Therefore, these models generated the errors ‘Model nearly unidentifiable’ and warnings about singular boundaries. These suggest we would not be able to reliably generate model estimates, so we do not report the results. We will change the manuscript to make this reasoning more explicit.

    Also, please standardize how the GLMM results are presented. There should be the estimate, SE, Z or t, then p value. (line 99, 137).

    We will update the results reporting to be standardised as the reviewer has suggested.

    Given the high rates of exposure to playbacks, I think the authors should test trial # (or video #) for a potential habituation effect, with earlier captures more likely to draw stronger responses than later video captures for each group.

    We will include video number in a reanalysis of the data.

    Reviewer #2 (Public review):

    Summary:

    The article describes an interesting methodology to test hypotheses about the impact of anthropogenic noise on a small arboreal primate, the pygmy marmoset. The authors used a motion-triggered combination of camera traps and speakers to play back control sounds, avian predator calls, and anthropogenic noise to test the risk-disturbance hypothesis and the distracted prey hypothesis. In addition, the authors implemented a technique that is usually used for larger mammals and has not been used before for smaller arboreal animals. The authors are careful in their interpretation of the results and do not favor one hypothesis over the other. The authors also elaborate extensively in their discussion on how to improve this kind of data collection in the future.

    Strengths:

    This study provides a method for rapid data collection while minimizing observer impact. The sample size is comparatively large for a wild animal in a reserve, given the overall observation time. The article also benefits from a solid analysis of the data.

    Weaknesses:

    Though the authors tested two contrasting hypotheses, the discussion would benefit from more detail on the ecological relevance of the observed behaviors.

    We thank the reviewer for their comments and we will update the discussion to add more detail on the ecological relevance of the behaviours that we observed.

    References

    Bui, S., Madaro, A., Nilsson, J., Fjelldal, P.G., Iversen, M.H., Brinchmann, M.F., Venås, B., Schrøder, M.B. and Stien, L.H. 2022. Warm water treatment increased mortality risk in salmon. Veterinary and Animal Science, 17, 100265. https://doi.org/10.1016/j.vas.2022.100265