Jul 28 (Tue) @ 4:00PM: "Co-Architecting Models and Silicon: Efficient Brain-Inspired Computing and Generative Visual Intelligence," Boxun Xu, ECE PhD Defense

Date and Time

Location: Harold Frank Hall, Room 4108 (Zoom Meeting: https://ucsb.zoom.us/j/8347381968)
Research Area: Computer Engineering
Research Keywords: Machine Learning / Neuromorphic Computing / AI, Computer Vision, Bio/Brain Inspired Computing, Computer Architecture / System Level Design

Abstract

Two of the most active frontiers in AI study intelligence from complementary directions, and both are bottlenecked by the same constraint: efficiency. One asks how intelligence arises, and builds brain-inspired spiking transformers whose computation is event-driven and highly sparse. The other asks how intelligence models the physical world, and builds generative visual models that synthesize images and video autoregressively. Each is powerful in principle yet unaffordable in practice: model and machine are designed in isolation, so neither meets the energy and latency budgets real workloads impose. 

This research argues that efficiency is not a post-hoc optimization, but a design principle spanning the full stack, from learning algorithms through model architectures, hardware architecture to the integrated circuits beneath. Its central mechanism is a single move repeated at every layer: identify the structure a workload already possesses, whether firing sparsity, scale locality, outlier geometry, or persistent attention, and make the algorithm, the architecture, and the circuits agree on it. One methodology makes both directions efficient.

The method reaches down both directions. For brain-inspired computing, it drives efficiency the whole way down the stack. In the model, spatiotemporal spiking attention and heterogeneous quantization expose the sparse structure that the hardware then exploits. On the hardware, it begins with SpikeX, an accelerator for sparse spiking convolutional networks, then builds Bishop, a dedicated accelerator and hardware-software co-design framework for spiking transformers, and rises into 3D-integrated hardware, extended to mixture-of-experts with dynamic head pruning. Visual generation is a separate problem, centered on visual autoregressive modeling and bound by memory and latency when serving. Here the method makes interpretation its instrument: beginning with AMS-KV, an adaptive multi-scale key-value cache, it reads why caches and attention behave as they do and turns that understanding into VAR-Q, an interpretable, tuning-free quantization framework, and Sparse Forcing, a native trainable sparse attention that brings autoregressive video generation into real time over longer horizons.

Co-designing the model and the computing platform beneath it, this work shows that efficient, real-time AI follows from one cross-layer discipline, bringing both ways of studying intelligence, how it arises and how it generates the visual world, within reach of the hardware that must run them.

Bio

Boxun Xu is a Ph.D. candidate in Electrical and Computer Engineering at UC Santa Barbara, advised by Prof. Peng Li. His research builds efficient and scalable artificial intelligence across the full stack, spanning algorithms, computer architecture, and integrated circuits, especially for brain-inspired computing and generative visual intelligence. His work reaches across machine learning, computer vision, computer architecture, and design automation, appearing at top venues including ISCA, AAAI, CVPR, ICCV, ICCAD, TCAD, ITC, and JSSC. The work he led was twice nominated for the ICCAD William J. McCalla Best Paper Award, in 2024 and 2025. He was a 2023 DAC Young Fellow and a recipient of the 2026 Radhakrishnan Nagarajan Family Fellowship, and received the ECE Outstanding Teaching Assistant Award in 2022 and 2023. His academic research has been complemented by research internships at Meta AI in 2024 and at Meta Superintelligence Labs in 2025 and 2026. Prior to UCSB, he received an M.S. from the University of Michigan, Ann Arbor, and a B.S. from the University of Electronic Science and Technology of China.

Hosted By: ECE Professor Peng Li

Submitted By: Boxun Xu <boxunxu@ucsb.edu>