Date of Award
6-26-2026
Date Published
August 2026
Degree Type
Dissertation
Degree Name
Doctor of Philosophy (PhD)
Department
Electrical Engineering and Computer Science
Advisor(s)
Qinru Qiu
Subject Categories
Computer Engineering | Engineering
Abstract
Deep neural networks have achieved state-of-the-art performance across a wide range of computer vision tasks, yet their deployment in real-world and safety-critical domains remains limited by two major challenges: vulnerability to adversarial perturbations and high computational cost. While adversarial training can improve robustness under certain threat models, it is computationally expensive and often generalizes poorly to black-box or transferable attacks. At the same time, conventional vision pipelines process uniformly sampled high-resolution images, introducing substantial data redundancy, latency, and energy overhead. In contrast, the human visual system achieves efficient and robust perception through space-variant foveal-peripheral sampling, attention-guided saccadic eye movements, and cortical filling-in mechanisms. Inspired by these principles, this dissertation develops a suite of bio-inspired active vision systems that reconstruct semantically coherent visual representations from sparse, localized glimpses selected by learned saccade policies. The proposed frameworks address efficiency and robustness across both conventional visual recognition models and large-scale vision-language architectures. First, we introduce a computational foveation framework for efficient visual sensing. The system integrates foveal-peripheral sampling, a multi-layer Convolutional Long Short-Term Memory predictive reconstruction network, and an Advantage Actor-Critic saccade controller. By actively steering high-resolution glimpses toward informative regions and reconstructing complete visual scenes from sparse observations, the framework reduces input pixel usage by up to 90\% while maintaining competitive classification accuracy. On ImageNet, it outperforms prior foveation-based baselines by 5\% in top-1 accuracy under comparable pixel budgets. Second, we extend this bio-inspired paradigm to adversarial robustness through SAFER-AiD: Saccade-Assisted Foveal-peripheral vision Enhanced Reconstruction for Adversarial Defense. SAFER-AiD serves as a classifier-agnostic, plug-and-play preprocessing module that combines sparse visual sampling, learned saccades, and predictive reconstruction. The sparse sampling process reduces the influence of localized adversarial noise, while the reconstruction module maps perturbed inputs toward semantically coherent visual representations. Under transferable attacks, including TGR and MI-FGSM, and gradient-free attacks such as SPSA, SAFER-AiD improves top-1 accuracy by up to 21.6\% across diverse CNN and Vision Transformer backbones without retraining downstream classifiers. Finally, we introduce Active Perception with EXploratory Sampling (APEX) to improve the robustness of large-scale vision-language models. APEX operates as a bio-inspired front-end for frozen CLIP architectures and is evaluated under adaptive adversarial attacks using a Straight-Through Estimator to approximate gradients through the saccade-based sampling process. We further combine APEX with a Label Vector Pool (LVP) to tackle robust Class-Incremental Learning (CIL). APEX achieves superior zero-shot and incremental robustness under strong $l_\infty$ perturbations, consistently suppressing adversarial drift while preserving preserving open-vocabulary generalization capability of the core foundation network. Together, these contributions demonstrate that biologically grounded active perception provides an effective interface for efficient and robust artificial vision. By combining sparse sensing, learned visual exploration, and predictive reconstruction, this dissertation offers a scalable pathway toward secure, energy-efficient, and edge-deployable deep learning systems.
Access
Open Access
Recommended Citation
Liu, Jiayang, "Bio-Inspired Visual Intelligence: Improving Efficiency and Robustness through Foveal-Peripheral Sampling and Learned Saccades" (2026). Dissertations - ALL. 2372.
https://surface.syr.edu/etd/2372
