AI in practice: a multilingual survey of 2025 BioHackathon participants
Curation statements for this article:-
Curated by GigaByte
Editors Assessment:
This Data Release presents a multilingual dataset collected from a survey of participants and community members at the 2025 DBCLS BioHackathon, with the goal of understanding how AIis being used in bioinformatics and related scientific fields. The survey, available in English, Japanese, and Thai, gathered responses from 105 participants about AI usage frequency, applications, challenges, institutional support, satisfaction, concerns, and demographic information. The released dataset includes anonymized raw responses, a cleaned English-language version for quantitative analysis, the original questionnaire, a data dictionary, and a translation lookup table to support reproducible research. To protect participant privacy, the authors carefully removed personally identifiable information. This openly available resource provides researchers with valuable data for studying AI adoption, evaluating policy and institutional support, and developing methods for analyzing survey data related to AI use in scientific research. Overall, this work contributes a well-documented, reusable dataset that can help inform future research on the evolving role of artificial intelligence in genomics, bioinformatics, software development, and the broader scientific community.
This evaluation refers to version 1 of the preprint
This article has been Reviewed by the following groups
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
- Endorsed by GigaByte (scotted400)
- Evaluated articles (GigaByte)
Abstract
This dataset arises from a multilingual survey of AI use among participants and community members in the DBCLS BioHackathon 2025 in Japan. The questionnaire, offered in English, Japanese, and Thai, asked about how often respondents use AI tools, what they use them for, obstacles they encounter, institutional support, satisfaction, and concerns. Additional items captured role, institution type, work country, and other demographics, totaling 105 responses. The dataset includes both raw anonymized responses and a cleaned, standardized English-only version suitable for quantitative analysis, along with the full questionnaire, a data dictionary for cleaned dataset, and a translation lookup table. Free-text answers were screened and redacted to remove URLs, names, and other potentially identifiable information. Together, these materials provide a community-level view of AI practice in genomics, bioinformatics, software development, and related areas, and can support work on AI adoption, policy, and methods for analyzing survey data on AI use in science.
Article activity feed
-
Editors Assessment:
This Data Release presents a multilingual dataset collected from a survey of participants and community members at the 2025 DBCLS BioHackathon, with the goal of understanding how AIis being used in bioinformatics and related scientific fields. The survey, available in English, Japanese, and Thai, gathered responses from 105 participants about AI usage frequency, applications, challenges, institutional support, satisfaction, concerns, and demographic information. The released dataset includes anonymized raw responses, a cleaned English-language version for quantitative analysis, the original questionnaire, a data dictionary, and a translation lookup table to support reproducible research. To protect participant privacy, the authors carefully removed personally identifiable information. This openly available resource …
Editors Assessment:
This Data Release presents a multilingual dataset collected from a survey of participants and community members at the 2025 DBCLS BioHackathon, with the goal of understanding how AIis being used in bioinformatics and related scientific fields. The survey, available in English, Japanese, and Thai, gathered responses from 105 participants about AI usage frequency, applications, challenges, institutional support, satisfaction, concerns, and demographic information. The released dataset includes anonymized raw responses, a cleaned English-language version for quantitative analysis, the original questionnaire, a data dictionary, and a translation lookup table to support reproducible research. To protect participant privacy, the authors carefully removed personally identifiable information. This openly available resource provides researchers with valuable data for studying AI adoption, evaluating policy and institutional support, and developing methods for analyzing survey data related to AI use in scientific research. Overall, this work contributes a well-documented, reusable dataset that can help inform future research on the evolving role of artificial intelligence in genomics, bioinformatics, software development, and the broader scientific community.
This evaluation refers to version 1 of the preprint
-
AbstractThis dataset arises from a multilingual survey of AI use among participants and community members in the DBCLS BioHackathon 2025 in Japan. The questionnaire, offered in English, Japanese, and Thai, asked about how often respondents use AI tools, what they use them for, obstacles they encounter, institutional support, satisfaction, and concerns. Additional items captured role, institution type, work country, and other demographics, totaling 105 responses. The dataset includes both raw anonymized responses and a cleaned, standardized English-only version suitable for quantitative analysis, along with the full questionnaire, a data dictionary for cleaned dataset, and a translation lookup table. Free-text answers were screened and redacted to remove URLs, names, and other potentially identifiable information. Together, these …
AbstractThis dataset arises from a multilingual survey of AI use among participants and community members in the DBCLS BioHackathon 2025 in Japan. The questionnaire, offered in English, Japanese, and Thai, asked about how often respondents use AI tools, what they use them for, obstacles they encounter, institutional support, satisfaction, and concerns. Additional items captured role, institution type, work country, and other demographics, totaling 105 responses. The dataset includes both raw anonymized responses and a cleaned, standardized English-only version suitable for quantitative analysis, along with the full questionnaire, a data dictionary for cleaned dataset, and a translation lookup table. Free-text answers were screened and redacted to remove URLs, names, and other potentially identifiable information. Together, these materials provide a community-level view of AI practice in genomics, bioinformatics, software development, and related areas, and can support work on AI adoption, policy, and methods for analyzing survey data on AI use in science.
This final version of this paper is published in GigaByte (see: https://doi.org/10.46471/gigabyte.179) and the peer reviews are available under a CC-BY license.
Reviewer 1. Gabrielle O'Brien
Are all data available and do they match the descriptions in the paper? No No. I don't actually see the dataset provided here, only the PDF of the article. It doesn't appear to be in the OSF repository either (perhaps I am looking in the wrong place? I have never reviewed for GigaByte before so it is possible). Are the data and metadata consistent with relevant minimum information or reporting standards? No. I don't actually see a copy of the data. I would like to see an exact list of the survey questions. It is not sufficient to report the topic they are addressing, the wording is important to have. Please attach a copy of the survey with the exact question wording and wording of response items. Is there sufficient detail in the methods and data-processing steps to allow reproduction? No. As commented above, there is insufficient information to reproduce the data collection method because the exact survey questions and response options are not provided (unless I am simply overlooking them?) Is there sufficient data validation and statistical analyses of data quality? No. I don't see any statistical analyses of data quality here. It looks like the data has been manually reviewed for quality, which may suffice depending on its size and complexity. But it would be nice to see some counts for completeness. Is the validation suitable for this type of data? Yes. I'm not entirely sure because I don't appear able to see the data itself, but I lean towards yes for the validations that are presented. Is there sufficient information for others to reuse this dataset or integrate it with other data? No As mentioned above, exact list of survey questions and response options is required and I don't see it.
Reviewer 2. Pichaya Lertvilai
Is the language of sufficient quality? Yes. The manuscript is well-written and the dataset is well-organized for a Data Release publication. The authors have taken appropriate steps toward anonymization and have structured the data for reuse.
Is the data acquisition clear, complete and methodologically sound? Yes. Comments Several categorical variables in the cleaned dataset contain free-text "Other" responses that fall outside the defined response categories. For example: - AI usage level: One response reads
What "AI" are you talking about exactly?! real Artificial-Intelligence or the trendy Artificial-Idiot (TM) !?and another readsit is complicated. These are non-standard responses to what the codebook describes as an "ordinal, single choice" variable. - Institution type: Entries includeunemployedandSmall business owner, which are outside the defined categories (Academia, Private sector, Public sector). The manuscript does not discuss how these "Other: free-text" responses should be handled in analysis. The codebook mentions that some variables include "Other: free-text" but does not enumerate the free-text values that appear. The authors should either (a) add a note in the codebook or manuscript clarifying how "Other" responses are represented, or (b) consider recoding these in the cleaned dataset.Is there sufficient detail in the methods and data-processing steps to allow reproduction? Yes
- Two column headers in the cleaned data CSV (and both raw data CSVs) contain the typo "Select all the apply" instead of "Select all that apply": - Column 2:
What is your field? (Select all the apply)— codebook has(Select all that apply)- Column 19:What would AI tools need to improve for you to be more satisfied? (Select all the apply)— codebook has(Select all that apply)This creates a mismatch between the codebook variable names and the actual CSV headers, which will cause problems for users attempting programmatic joins between the codebook and data files. These typos should be corrected to ensure consistency, or the codebook should match the exact header strings used in the data. - The codebook contains 16 entries, but the cleaned dataset has 26 columns. This is because two multi-part questions are represented as single codebook entries: - The "task usage" question expands into 7 sub-columns (Coding, Research, Brainstorming, Writing/Editing, Teaching and curriculum, Translation, Personal use) - The "AI concern" question expands into 5 sub-columns (Bias in algorithms, Data privacy/security, Intellectual property/ownership, Misinformation/Hallucinations, Environment impact) While the codebook mentions the categories, it does not list the 26 individual column names that appear in the CSV. For a dataset intended for reuse, the codebook should ideally have one entry per column (i.e., 26 entries), listing the exact column header as it appears in the CSV file. This would prevent confusion and support automated data dictionary workflows.
Is there sufficient data validation and statistical analyses of data quality? Yes.
- The cleaned dataset has substantial missing data in several columns that goes undiscussed in the manuscript: - "Is there a reason why you don't use AI?": 99% missing (expected, as this applies only to non-users) - "What would AI tools need to improve?": 68% missing - "Teaching and curriculum" task usage: 41% missing - "Data privacy/security" concern: 39% missing - "Intellectual property/ownership" concern: 33% missing - "Bias in algorithms" concern: 28% - "Environment impact" concern: 28% While the manuscript notes that "almost all questions were optional," the high missing rates for specific items should be mentioned in the Data Validation section, especially if the pattern is non-random (e.g., respondents who skip concern items may differ systematically from those who answer). This information is important for re-users to make informed analytical choices.
- All three data files (raw original, raw English, cleaned English) contain exactly 105 data rows, confirming structural consistency. However, the manuscript does not report how many total survey starts occurred versus completions, nor whether any responses were excluded (e.g., duplicate submissions, empty submissions). A brief statement on this would strengthen the data validation section.
- The manuscript states (Data Validation section): "One response containing an occupation ('Postdoctoral researcher') was set to missing." This correction was properly applied in the cleaned English file (row 70 is blank), but the raw English file (
Biohackathon2025AISurvey_data_raw_anon_ENG_v1.csv, row 70) still contains "Postdoctoral researcher" in the country column, as does the original-languages raw file. The manuscript should clarify whether the raw files are intentionally left uncorrected (preserving the original response as-is), or whether this correction was inadvertently omitted from the raw files. If the former, this design decision should be explicitly stated.
Is there sufficient information for others to reuse this dataset or integrate it with other data? Yes. The manuscript acknowledges that "minor spelling and formatting variants were left as entered." However, in the cleaned English dataset (described as "suitable for quantitative analysis"), country values remain unstandardized. For example: - "germany" (lowercase) vs. "Germany" (9 responses) - "U.K." (1 response) vs. "United Kingdom" (2 responses) - "USA" (2 responses) vs. "United States" (1 response) - "Global" (1 response) — an ambiguous non-country value For a dataset intended for quantitative cross-tabulation, these inconsistencies will require downstream users to perform their own harmonization. The authors should consider either standardizing country names in the cleaned dataset or documenting this as a known limitation for re-users Additional Comments: The translation lookup table (Translation_lookup_table_v1.csv) is showing incorrect unicode formatting for Thai and Japanese languages when users download the data set. It is recommended that the authors try to make sure that the format for other languages is well preserved during data upload/download.
- Two column headers in the cleaned data CSV (and both raw data CSVs) contain the typo "Select all the apply" instead of "Select all that apply": - Column 2:
-
-