O desempenho de grandes modelos de linguagem na síntese de relatos de pesquisa educacional: Estudo comparativo com artigos no idioma português
This article has been Reviewed by the following groups
Listed in
- Evaluated articles (PREreview)
Abstract
Este estudo objetiva comparar o desempenho do ChatGPT 5.0, Gemini 3.0 Flash e DeepSeek V3 na síntese de artigos científicos publicados em português. Adotou-se abordagem mista, com teste de geração de resumos para cinco artigos da área da educação em três variações de prompt. A análise dos outputs foi realizada por 15 especialistas nas temáticas dos artigos que utilizaram uma escala ordinal de Percepção de Qualidade (PQ) para mensurar as dimensões de qualidade de conteúdo e isenção de vieses. Os resultados indicaram bom desempenho dos modelos, com resultados de PQ entre 84% e 94% do escore total. O teste de Kruskal-Wallis identificou diferenças significativas na comparação dos resultados por artigo, (H(4) = 24,040, p < 0,001) e inexistência de diferenças entre os modelos (H(2) = 1,604, p = 0,448) ou tipos de prompt (H(8) = 2,506, p = 0,961), sugerindo paridade tecnológica na tarefa de síntese de textos científicos. Ao lado disso, a análise qualitativa identificou tendência à generalização e à supressão de marcadores de incerteza, evidências que indicam falhas na validade dos resumos produzidos. Os resultados obtidos também sugerem potenciais vieses oriundos do idioma de treinamento dos modelos, que, hipoteticamente, podem afetar a captura da textura narrativa e de jargão específico da ciência produzida em português. Concluiu-se que, embora os modelos de IAG testados apresentem bom desempenho sintático, a supervisão humana permanece indispensável para garantia da fidedignidade e integralidade da comunicação científica da pesquisa em educação.
Article activity feed
-
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23089735.
Peer Review: Epistemic Compression and Uncertainty Suppression in Generative AI Summarization
Preprint Reviewed: O desempenho de grandes modelos de linguagem na síntese de relatos de pesquisa educacional: Estudo comparativo com artigos no idioma português (Martins et al., SciELO Preprints, DOI: 10.1590/scielopreprints.17611)
Summary & Core Findings
The authors evaluate how leading artificial intelligence systems synthesize academic literature in Portuguese. Using an expert evaluation panel and non-parametric statistical testing, the study reveals significant performance parity across different AI tools and prompting strategies.
The most critical takeaway is qualitative: despite …
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23089735.
Peer Review: Epistemic Compression and Uncertainty Suppression in Generative AI Summarization
Preprint Reviewed: O desempenho de grandes modelos de linguagem na síntese de relatos de pesquisa educacional: Estudo comparativo com artigos no idioma português (Martins et al., SciELO Preprints, DOI: 10.1590/scielopreprints.17611)
Summary & Core Findings
The authors evaluate how leading artificial intelligence systems synthesize academic literature in Portuguese. Using an expert evaluation panel and non-parametric statistical testing, the study reveals significant performance parity across different AI tools and prompting strategies.
The most critical takeaway is qualitative: despite generating highly fluent, professional-sounding summaries, AI models systematically over-generalize text and actively strip out uncertainty markers, nuance, and conditional logic from the original research.
Core Strengths & Real-World Implications
Fluency as a Mask for Information Loss:
The finding that AI models remove context and conditional statements highlights a major operational danger. When automated tools generate clean, confident summaries while stripping away underlying risks or limitations, they create a false sense of completeness. This artificial polish lulls human reviewers into trusting incomplete outputs—a dynamic I define in my research on Systemic Disclosure Architecture as blind algorithmic deference.
Structural Limitations Over Simple Prompting:
Because changing the prompt structure failed to stop models from flattening narrative nuance, this distortion is clearly a fundamental architectural limitation of generative text models, not a user formatting error. In my work on the Systemic Intent Shadow, I frame this as an epistemic gap: automated systems produce confident, smooth representations that detach entirely from the complex, messy realities of the source material.
Direct Recommendations for the Authors
Measure What Was Erased:
Track hedges, conditional statements, and confidence bounds before and after summarization. Adding a dedicated metric for how much uncertainty was preserved will prove far more useful than measuring surface readability alone.
Warn Readers About High Readability Scores:
State clearly in the conclusion that high surface quality acts as a trap. When an AI summary looks perfect on the surface, human operators lower their guard—making human oversight mandatory in academic, legal, and institutional settings.
Conclusion
A solid, highly practical study. It clearly proves that surface fluency in AI routinely conceals critical data loss. Highly recommended for formal publication.
Reviewer:
Julián Rodríguez, Jr., FRSA, MRES, M.ISRM
Managing Principal, Julian Rodriguez & Associates
ORCID: 0009-0007-9332-0140
Related Frameworks for Further Reading:
Rodriguez, J., Jr. (2026). The Illusion of Completeness: Systemic Disclosure Architecture and the Mitigation of Algorithmic Deference in Institutional Environments. Zenodo. https://doi.org/10.5281/zenodo.23022219
Rodriguez, J., Jr. (2026). The Systemic Intent Shadow: Mapping Organizational Collapse and Regulatory Friction in the Era of AI-Driven Governance. Zenodo. https://doi.org/10.5281/zenodo.22304608
Competing interests
The author declares that they have no competing interests.
Use of Artificial Intelligence (AI)
The author declares that they used generative AI to come up with new ideas for their review.
-