Design of domain-specific FPGA architectures for accelerating AI/ML inference, signal processing, and communication workloads, with an emphasis on high performance, low latency, and energy efficiency.
Custom RISC-V processors and SoC architectures incorporating specialised instructions, tightly coupled accelerators, and efficient processor–accelerator integration.
Hardware optimisation through high-level synthesis, parallelism, pipelining, dataflow and memory-system design, quantised arithmetic, and performance–area–energy trade-offs.
Integrated algorithm and architecture design for efficient AI deployment on edge and heterogeneous platforms, including emerging intelligent and semantic communication applications.