O desempenho de grandes modelos de linguagem na síntese de relatos de pesquisa educacional: Estudo comparativo com artigos no idioma português

This article has been Reviewed by the following groups

Read the full article

Abstract

Este estudo objetiva comparar o desempenho do ChatGPT 5.0, Gemini 3.0 Flash e DeepSeek V3 na síntese de artigos científicos publicados em português. Adotou-se abordagem mista, com teste de geração de resumos para cinco artigos da área da educação em três variações de prompt. A análise dos outputs foi realizada por 15 especialistas nas temáticas dos artigos que utilizaram uma escala ordinal de Percepção de Qualidade (PQ) para mensurar as dimensões de qualidade de conteúdo e isenção de vieses. Os resultados indicaram bom desempenho dos modelos, com resultados de PQ entre 84% e 94% do escore total. O teste de Kruskal-Wallis identificou diferenças significativas na comparação dos resultados por artigo, (H(4) = 24,040, p < 0,001) e inexistência de diferenças entre os modelos (H(2) = 1,604, p = 0,448) ou tipos de prompt (H(8) = 2,506, p = 0,961), sugerindo paridade tecnológica na tarefa de síntese de textos científicos. Ao lado disso, a análise qualitativa identificou tendência à generalização e à supressão de marcadores de incerteza, evidências que indicam falhas na validade dos resumos produzidos. Os resultados obtidos também sugerem potenciais vieses oriundos do idioma de treinamento dos modelos, que, hipoteticamente, podem afetar a captura da textura narrativa e de jargão específico da ciência produzida em português. Concluiu-se que, embora os modelos de IAG testados apresentem bom desempenho sintático, a supervisão humana permanece indispensável para garantia da fidedignidade e integralidade da comunicação científica da pesquisa em educação.

Article activity feed

  1. This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23089735.

    Peer Review: Epistemic Compression and Uncertainty Suppression in Generative AI Summarization

    Preprint Reviewed: O desempenho de grandes modelos de linguagem na síntese de relatos de pesquisa educacional: Estudo comparativo com artigos no idioma português (Martins et al., SciELO Preprints, DOI: 10.1590/scielopreprints.17611)

    Summary & Core Findings

    The authors evaluate how leading artificial intelligence systems synthesize academic literature in Portuguese. Using an expert evaluation panel and non-parametric statistical testing, the study reveals significant performance parity across different AI tools and prompting strategies.

    The most critical takeaway is qualitative: despite generating highly fluent, professional-sounding summaries, AI models systematically over-generalize text and actively strip out uncertainty markers, nuance, and conditional logic from the original research.

    Core Strengths & Real-World Implications

    • Fluency as a Mask for Information Loss:

      The finding that AI models remove context and conditional statements highlights a major operational danger. When automated tools generate clean, confident summaries while stripping away underlying risks or limitations, they create a false sense of completeness. This artificial polish lulls human reviewers into trusting incomplete outputs—a dynamic I define in my research on Systemic Disclosure Architecture as blind algorithmic deference.

    • Structural Limitations Over Simple Prompting:

      Because changing the prompt structure failed to stop models from flattening narrative nuance, this distortion is clearly a fundamental architectural limitation of generative text models, not a user formatting error. In my work on the Systemic Intent Shadow, I frame this as an epistemic gap: automated systems produce confident, smooth representations that detach entirely from the complex, messy realities of the source material.

    Direct Recommendations for the Authors

    • Measure What Was Erased:

      Track hedges, conditional statements, and confidence bounds before and after summarization. Adding a dedicated metric for how much uncertainty was preserved will prove far more useful than measuring surface readability alone.

    • Warn Readers About High Readability Scores:

      State clearly in the conclusion that high surface quality acts as a trap. When an AI summary looks perfect on the surface, human operators lower their guard—making human oversight mandatory in academic, legal, and institutional settings.

    Conclusion

    A solid, highly practical study. It clearly proves that surface fluency in AI routinely conceals critical data loss. Highly recommended for formal publication.

    Reviewer:

    Julián Rodríguez, Jr., FRSA, MRES, M.ISRM

    Managing Principal, Julian Rodriguez & Associates

    ORCID: 0009-0007-9332-0140

    Related Frameworks for Further Reading:

    • Rodriguez, J., Jr. (2026). The Illusion of Completeness: Systemic Disclosure Architecture and the Mitigation of Algorithmic Deference in Institutional Environments. Zenodo. https://doi.org/10.5281/zenodo.23022219

    • Rodriguez, J., Jr. (2026). The Systemic Intent Shadow: Mapping Organizational Collapse and Regulatory Friction in the Era of AI-Driven Governance. Zenodo. https://doi.org/10.5281/zenodo.22304608

    Competing interests

    The author declares that they have no competing interests.

    Use of Artificial Intelligence (AI)

    The author declares that they used generative AI to come up with new ideas for their review.