面向资源受限设备的XNOR-Popcount高能效边缘检测硬件模块
关键词:
近似计算; 边缘检测; 低功耗硬件设计; 乘累加器; XNOR-Popcount摘要
边缘检测是无人机导航、物联网摄像头与可穿戴设备等众多嵌入式视觉任务的基础构件。然而,基于乘累加(MAC)运算的传统边缘检测器难以适应此类资源受限硬件在功耗与面积上的严苛约束。本文提出一种完全可综合的Prewitt边缘检测器,以1比特XNOR–Popcount逻辑取代MAC运算。输入的8位像素与±1卷积核系数先经二值化,再由并行XNOR门处理,并通过轻量级Popcount加法树完成计数,从而彻底消除全部乘法器与DSP单元。在Xilinx Zynq-7020 FPGA上完成原型验证的结果表明,所提设计使查找表用量减少55%,触发器数量减少26%,动态功耗降低约60%,且可支持的时钟频率最高可达MAC型核心的5倍。在MNIST与ORL数据集上的帧级评估显示,其边缘保真度接近无损,单幅图像的差异度得分低于0.08,吞吐量提升接近4倍。上述结果表明,面向硬件的二值近似能够在不牺牲功能精度的前提下,为嵌入式人工智能系统提供实时、高能效的边缘检测能力。Abstract
Edge detection is a fundamental building block in many embedded vision tasks, including drone navigation, IoT cameras, and wearable devices. However, traditional edge detectors based on multiply–accumulate (MAC) operations are poorly suited to the tight power and area budgets of such resource-constrained hardware. This work introduces a fully synthesizable Prewitt edge detector that replaces MAC operations with 1-bit XNOR–Popcount logic. Incoming 8-bit pixels and ±1 kernel coefficients are binarized, processed by parallel XNOR gates, and tallied by a lightweight Popcount adder tree, eliminating all multipliers and DSP slices. Prototyped on a Xilinx Zynq-7020 FPGA, the proposed design reduces lookup-table usage by 55% and flip-flop count by 26%, cuts dynamic power by about 60%, and supports clock frequencies up to five times higher than a MAC-based core. Frame-level evaluations on the MNIST and ORL datasets show near-lossless edge fidelity, with per-image dissimilarity scores below 0.08 and throughput gains approaching four times. These results demonstrate that hardware-aware binary approximations can enable real-time, energy-efficient edge detection for embedded AI systems without sacrificing functional accuracy.References
[1] 周创兵. 岩质边坡稳定性智能监测方法. 岩土力学, 2024, 43(03): 599-608.
[2] 刘吉臻. 新能源火电耦合调峰系统. 动力工程学报, 2022, 42(04): 265-273.
[3] Sze V, Chen Y H, Yang T J, et al. Efficient processing of deep neural networks: a tutorial and survey. Proceedings of the IEEE, 2017, 105(12): 2295-2329.
[4] Redmon J, Farhadi A. YOLOv3: an incremental improvement. arXiv preprint, 2018.
[5] 周秉根. 风电叶片复合材料成型工艺. 复合材料学报, 2022, 39(03): 1067-1078.
[6] Burrello A, Garofalo A, Bruschi N, et al. DORY: automatic end-to-end deployment of real-world DNNs on low-cost IoT MCUs. IEEE Transactions on Computers, 2021, 70(8): 1253-1268.
[7] Horowitz M. Computing's energy problem (and what we can do about it). Digest of Technical Papers, IEEE International Solid-State Circuits Conference, 2014: 10-14.
[8] 陆建华. 6G通信超宽带传输技术. 电子学报, 2024, 50(03): 513-522.
[9] 李爱群. 老旧建筑抗震加固改造技术. 工程力学, 2022, 39(04): 1-10.
[10] Jouppi N P, et al. In-datacenter performance analysis of a tensor processing unit. Proceedings of the International Symposium on Computer Architecture, 2017: 1-12.
[11] Chen Y H, Krishna T, Emer J S, et al. Eyeriss: an energy-efficient reconfigurable accelerator for deep convolutional neural networks. IEEE Journal of Solid-State Circuits, 2017, 52(1): 127-138.
[12] Han S, Mao H, Dally W J. Deep compression: compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv preprint, 2015.