Mapping ML Workloads to Silicon: A Systematic Analysis of Accelerator Architectures and a Phase-Adaptive, SRAM-First Design (FMCT, codename Habanero)
Description
Machine learning's explosive growth has produced a proliferation of accelerator architectures, yet no framework maps a workload's computational character to the hardware that serves it best. We build that framework: decomposing ML workloads — transformers, CNNs, GNNs, diffusion and state-space models — into core compute primitives characterized by arithmetic intensity, memory-access pattern, and parallelism, and evaluating how GPUs, TPUs, systolic arrays, dataflow processors, wafer-scale engines, and neuromorphic chips serve them. Using hierarchical roofline and utilization modelling, we show the binding constraint has shifted from peak FLOPS to memory bandwidth and data-movement energy — which quantization only sharpens — and consolidate eight persistent, unsolved problems.
These motivate the central contribution: the Fused Memory-Compute Tile (FMCT, codename Habanero), an SRAM-first, HBM-free accelerator combining a three-mode (GEMM / bandwidth / fused) phase-adaptive tile switched by compiler directive, a non-linear unit fused into the systolic datapath, compute-in-memory as a selectable mode rather than the whole architecture, and an inline KV-cache compression engine. A validated component-level energy model shows ~4-7× batch-1 efficiency gains over HBM GPUs. We position FMCT against its closest SRAM-first, no-HBM analogs — d-Matrix Corsair and Tenstorrent — and OpenAI's HBM-retaining Jalapeño ASIC, closing with principles for memory-centric, workload-adaptive accelerators.
Files
AI_Chip_Architecture.pdf
Additional details
References
- Jouppi, N.P. et al. (2017). "In-Datacenter Performance Analysis of a Tensor Processing Unit." *ISCA '17*. DOI: 10.1145/3079856.3080246
- Chen, Y.-H. et al. (2017). "Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Neural Networks." *IEEE JSSC*, 52(1). DOI: 10.1109/JSSC.2016.2616357
- Sze, V. et al. (2017). "Efficient Processing of Deep Neural Networks: A Tutorial and Survey." *Proc. IEEE*, 105(12). DOI: 10.1109/JPROC.2017.2761740
- Williams, S. et al. (2009). "Roofline: An Insightful Visual Performance Model for Multicore Architectures." *CACM*, 52(4). DOI: 10.1145/1498765.1498785
- Dao, T. et al. (2022). "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness." *NeurIPS 2022*. arXiv: 2205.14135
- Dettmers, T. et al. (2022). "GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale." *NeurIPS 2022*. arXiv: 2208.07339
- Kwon, H. et al. (2019). "Understanding Reuse, Performance, and Hardware Cost of DNN Dataflows." *MICRO 2019*. DOI: 10.1145/3352460.3358252
- Narayanan, D. et al. (2021). "Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM." *SC '21*. arXiv: 2104.04473
- Shao, Y.S. et al. (2019). "Simba: Scaling Deep-Learning Inference with Multi-Chip-Module-Based Architecture." *MICRO 2019*. DOI: 10.1145/3352460.3358302
- Norrie, T. et al. (2021). "The Design Process for Google's Training Chips: TPUv2 and TPUv3." *IEEE Micro*, 41(2). DOI: 10.1109/MM.2021.3058217
- Lie, S. (2023). "Cerebras Architecture Deep Dive." *IEEE Micro*, 43(3). DOI: 10.1109/MM.2023.3262110
- Abts, D. et al. (2022). "A Software-Defined Tensor Streaming Multiprocessor for Large-Scale Machine Learning." *ISCA 2022*. DOI: 10.1145/3470496.3527405
- Dao, T. (2023). "FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning." arXiv: 2307.08691
- Frantar, E. et al. (2023). "GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers." *ICLR 2023*. arXiv: 2210.17323
- Dally, W.J., Turakhia, Y. & Han, S. (2020). "Domain-Specific Hardware Accelerators." *CACM*, 63(7), 48-57. DOI: 10.1145/3361682
- NVIDIA (2024). "NVIDIA Blackwell Architecture Technical Brief." https://resources.nvidia.com/en-us-blackwell-architecture
- Ivanov, A. et al. (2021). "Data Movement Is All You Need: A Case Study on Optimizing Transformers." *MLSys 2021*. arXiv: 2007.00072
- Reuther, A. et al. (2022). "AI and ML Accelerator Survey and Trends." *HPEC 2022*. arXiv: 2210.04055. DOI: 10.1109/HPEC55821.2022.9926331
- Kung, H.T. (1982). "Why Systolic Architectures?" *Computer*, 15(1). DOI: 10.1109/MC.1982.1653825
- Hennessy, J.L. & Patterson, D.A. (2019). "A New Golden Age for Computer Architecture." *CACM*, 62(2). DOI: 10.1145/3282307
- Vaswani, A. et al. (2017). "Attention Is All You Need." *NeurIPS 2017*. arXiv: 1706.03762
- Khwa, W.-S. et al. (2025). "A mixed-precision memristor and SRAM compute-in-memory AI processor." *Nature*, 639(8055), 617-623. DOI: 10.1038/s41586-025-08639-2
- Hua, S. et al. (2025). "An integrated large-scale photonic accelerator with ultralow latency." *Nature*, 640, 361-367. DOI: 10.1038/s41586-025-08786-6
- Tsirigotis, A. et al. (2025). "Photonic neuromorphic accelerator for convolutional neural networks." *Communications Engineering*, 4, Article 80. DOI: 10.1038/s44172-025-00416-3
- Lam, S. et al. (2026). "Neuromorphic photonic computing with electro-optic analog memory." *Nature Communications*. DOI: 10.1038/s41467-026-69084-x. arXiv: 2401.16515
- Shah, J. et al. (2024). "FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision." *NeurIPS 2024*. arXiv: 2407.08608
- UALink Consortium (2025). "Ultra Accelerator Link (UALink) Specification 1.0."
- JEDEC (2025). "JESD270-4 High Bandwidth Memory (HBM4) Standard." April 16, 2025.
- Jia, Z. et al. (2019). "Dissecting the Graphcore IPU Architecture via Microbenchmarking." arXiv: 1912.03413. *(Also: Knowles, S. (2021). "Graphcore Colossus Mk2 IPU." Hot Chips 33.)*
- Talpes, E. et al. (2023). "Tesla Dojo: Scaling Custom Silicon for ML Training." *IEEE Micro*, 43(5). DOI: 10.1109/MM.2023.3311422
- Lee, J. et al. (2024). "MTIA: First Generation Silicon Targeting Meta's Recommendation Systems." *ISCA 2024*. DOI: 10.1109/ISCA59077.2024.00023
- Reddi, V.J. et al. (2020). "MLPerf Inference Benchmark." *ISCA 2020*. DOI: 10.1109/ISCA45697.2020.00045
- Ahn, J. et al. (2015). "A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing." *ISCA 2015*. DOI: 10.1145/2749469.2750386
- Gao, M. et al. (2017). "TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory." *ASPLOS 2017*. DOI: 10.1145/3037697.3037702
- Kwon, H. et al. (2018). "MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects." *ASPLOS 2018*. DOI: 10.1145/3173162.3173176
- Muralimanohar, N., Balasubramonian, R. & Jouppi, N.P. (2009). "CACTI 6.0: A Tool to Model Large Caches." *HP Labs Technical Report HPL-2009-85*. *(Used with 5nm scaling projections from Stillmaker & Baas (2017), DOI: 10.1109/JSSC.2017.2680923)*
- Roune, B.H. (2026). "Designing AI Chip Software and Hardware." Unpublished technical document. https://docs.google.com/document/d/1dZ3vF8GE8_gx6tl52sOaUVEPq0ybmai1xvu3uk89_is *(Former TPUv3 software lead; proposes AI CPU architecture with heterogeneous systolic arrays, Int7+1 sparse numerics, hardware Huffman compression, managed aggregation, and C++ software pipelining library.)*
- Liu, Z. et al. (2024). "KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache." arXiv:2402.02750. *(Per-channel Key quantization + per-token Value quantization; 2-bit KV with <0.1% accuracy loss on Llama-2.)*
- Hooper, C. et al. (2024). "KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization." arXiv:2401.18079. *(Per-channel quantization with outlier-aware calibration; enables 1M+ context on single GPU.)*
- Kang, Y. et al. (2024). "Gear: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM." arXiv:2403.05527. *(Quantization + sparse outlier coding; 2-4× compression with negligible perplexity impact.)*
- Ma, S. et al. (2024). "The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits." arXiv:2402.17764. *(Demonstrates ternary {-1, 0, +1} weight models match FP16 accuracy from 3B params; eliminates multiplication, 55-82% energy reduction per token.)*
- Wang, J. et al. (2025). "BitNet b1.58 2B4T Technical Report." arXiv:2504.12285. *(First open-source natively trained 1.58-bit LLM at 2B params on 4T tokens; matches INT4-quantized models of similar size.)*
- d-Matrix (2025). "Corsair: Digital In-Memory Compute (DIMC) AI Inference Platform." Hot Chips 2025. https://www.d-matrix.ai/product/ *(Production inference accelerator: ~2 GB SRAM across chiplets, LPDDR5X capacity memory and no HBM, PCIe-5 card, TSMC 6nm, ~38 TOPS/W; ~30K tok/s at 2 ms/token on Llama3-70B. Closest commercial analog to the FMCT SRAM-first + digital-CIM + LPDDR5 + chiplet stack.)*
- Tenstorrent (2024–2026). "Blackhole and Quasar: RISC-V Tensix tile-mesh AI processors." https://tenstorrent.com *(Mesh of Tensix cores, each with multiple RISC-V cores + local SRAM and a dual-NoC mesh; GDDR6 and no HBM; stackable chiplet (Quasar) and a fully open RISC-V software stack. Embodies the RISC-V tile-mesh, no-HBM, chiplet and open-ISA arguments.)*
- Etched (2024–2026). "Sohu: Transformer-Only Inference ASIC." https://www.etched.com *(Transformer-only ASIC, TSMC 4nm, 144 GB HBM3E/chip, ~90% FLOP utilization, ~500K tok/s Llama-70B on an 8-chip server; cannot run CNN/LSTM/SSM. Extreme-specialization data point for the GEMM-fixation argument.)*
- Google (2025). "TPU v7 (Ironwood)." Google Cloud documentation. https://cloud.google.com/tpu *(7th-generation, inference-focused TPU: 4,614 FP8 TFLOPS, 192 GB HBM3E, 7.37 TB/s; pods to 9,216 chips, 42.5 FP8 ExaFLOPS.)*
- Microsoft (2026). "Maia 200 AI Inference Accelerator." https://blogs.microsoft.com/blog/2026/01/26/maia-200 *(TSMC 3nm, native FP8/FP4, 216 GB HBM3e at 7 TB/s, 272 MB on-chip SRAM, dedicated data-movement engines; inference-optimized.)*
- Amazon Web Services (2025). "AWS Trainium3." re:Invent 2025. https://aws.amazon.com/ai/machine-learning/trainium/ *(2.52 PFLOPS FP8 per chip; ~4.4× compute and ~40% better energy efficiency vs Trainium2.)*
- Kim, S. et al. (2024). "Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference." arXiv:2406.08413. *(Survey of analog and digital CIM (SRAM/NVM/DRAM) for LLM inference; anchor reference for the CIM landscape.)*
- Houshmand, P. et al. (2023). "Benchmarking and Modeling of Analog and Digital SRAM In-Memory Computing Architectures." arXiv:2305.18335. *(Quantitative analog-vs-digital SRAM-CIM comparison; supports the choice of digital CIM at 4–8 bit precision.)*
- NVIDIA (2026). "Vera Rubin Platform — Six New Chips, One AI Supercomputer." CES/GTC 2026. https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer *(Next-gen platform: Rubin R100 GPU with 288 GB HBM4 at 22 TB/s and ~50 PFLOPS NVFP4 inference, plus the Vera CPU, NVLink 6 switch, ConnectX-9, BlueField-4, and Spectrum-6; ~5× Blackwell inference. Includes the Rubin CPX "context" GPU dedicated to long-context prefill.)*
- AMD (2026). "Instinct MI400 / MI450 series (CDNA5) and Helios rack." CES 2026. https://www.amd.com/en/products/accelerators/instinct.html *(MI450: 432 GB HBM4 at 19.6 TB/s, ~40 PFLOPS FP4 / 20 PFLOPS FP8; MI430X/MI440X/MI455X variants; competes with Vera Rubin on memory capacity and scale-out.)*
- Qualcomm (2025). "AI200 and AI250 data-center inference accelerators." https://www.qualcomm.com/products/technology/processors/ai-accelerators *(AI200 (2026): 768 GB LPDDR5 per card — total on-card capacity, not HBM bandwidth, as the differentiator; AI250 (2027). A capacity-first, LPDDR-based, no-HBM point in the same cluster as FMCT and d-Matrix.)*
- NVIDIA (2026). "RTX 50-series (Blackwell) consumer GPUs." https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/ *(RTX 5090: 32 GB GDDR7, 1,792 GB/s, 575 W, \$1,999; 5080/5070 Ti/5070 complete the consumer lineup. Reference point for the consumer-GPU comparison in §14.13.)*
- OpenAI / Broadcom (2026). "Jalapeño: LLM-optimized inference processor." https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ ; https://openai.com/index/openai-and-broadcom-announce-strategic-collaboration/ *(OpenAI's first custom chip, announced 2026-06-24; collaboration signed Oct 2025. Inference-only ASIC on TSMC 3nm, ~840 mm² near the reticle limit; 2.5D interposer with one central compute die + 6–8 HBM3/HBM4 stacks + I/O chiplet + 2 dummy dies; systolic-array compute; scaled out over Broadcom Ethernet rather than NVLink. Claims "significantly better performance per watt" and — per Broadcom CEO Hock Tan, not OpenAI's own release — ~50% lower cost per token vs NVIDIA GPUs. 9-month design→tape-out; part of a 10 GW deployment program (2H2026→2029). The HBM-retaining inference-ASIC counterpoint to FMCT's HBM-free design; see §14.7.2 and §14.13.)*