Large-Model-Assisted Analysis of Adversarial Vulnerability: Separating Evidence-Conflict Detection from Robust Classification

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Adversarial examples reveal a persistent mismatch between human visual invari-ance and the decision boundaries learned by deep neural networks. We study this mismatch as a mechanism problem rather than as a search for a larger backbone or a state-of-the-art robustness number. The proposed workflow uses a large model as an analysis assistant to generate hypotheses about fragile evidence. It then tests these hypotheses with paired clean–adversarial tensors, evidence-switching gates, and white-box attacks. We instantiate the idea on CIFAR-10 with compact ResNet-style backbones, an Aha-ResNet gate, a pairwise ranking variant called AhaV2, and GRPO-style policies that select among base, low-pass, edge, and semantic views. Across tested remote GPU configurations, AhaV2 and GRPO induce larger trigger or semantic-action shifts on adversar-ial inputs than on clean inputs. The best long-run AhaV2 model reaches 79.85% clean accuracy, 37.45% FGSM accuracy, and 6.65% PGD-50 two-restart accuracy. Human-aligned and joint semantic variants increase PGD-50 Aha shifts to 27.80 and 27.40 percentage points, respectively, but do not improve strong white-box accuracy. A PGD-5 plug-and-play study further shows that a 1,495-parameter AhaV2 trigger improves clean and FGSM accuracy on a frozen adversarially trained backbone, while only slightly improving PGD-50 accuracy. On TinyIma-geNet, the proposed Aha lens separates PGD-AT, TRADES, MART, and AWP into robust-boundary improvement and unresolved evidence-conflict detection. The main finding is diagnostic: detecting adversarial conflict is trainable and measurable, but it is separable from possessing a robust decision boundary in the alternative evidence path.

Article activity feed