semantic-scholar
Semantic ScholarExtracts: paper title, abstract snippet · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public endpoint; no project credential required.
Last ingested: 2026-08-12 · run status: on demand
Archive source — full history has value. Use pagination to browse older records.
Query: photonic computing transformer accelerator Authors: S. Serunjogi, Rachmad Vidya Wicaksana Putra, Ayat M. Taha, Muhammad Shafique, Mahmoud S. Rasras Citations: 0 Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy eff
Query: ai datacenter power cooling efficiency Authors: Wedan Emmanuel Gnibga, A. Chien Citations: 0 Datacenter (DC) power demand is rapidly growing to support AI workloads. With increasing pressure to reduce water consumption, DC cooling power can be up to 30% overall DC power, r
Query: mixture of experts inference serving Authors: Craig Opie Citations: 0 Sparse Mixture-of-Experts (MoE) language models separate total parameter count from per-token active computation, but local inference systems often still require the full model, key-value cache, runtime
Query: large language model inference gpu datacenter Authors: Sicheng Zhou, Rohan Basu Roy Citations: 0 Large language model (LLM) inference consumes substantial datacenter energy and produces both carbon emissions and criteria pollutants (PM2.5, SO2, NOx). These criteria polluta
Query: large language model inference gpu datacenter Authors: Sicheng Zhou, Rohan Basu Roy Citations: 0 Large language model (LLM) inference consumes substantial datacenter energy and produces both carbon emissions and criteria pollutants (PM2.5, SO2, NOx). These criteria polluta
Query: large language model inference gpu datacenter Authors: Sicheng Zhou, Rohan Basu Roy Citations: 0 Large language model (LLM) inference consumes substantial datacenter energy and produces both carbon emissions and criteria pollutants (PM2.5, SO2, NOx). These criteria polluta
Query: high bandwidth memory artificial intelligence accelerator Authors: Sing-Da Jiang, Jing-An Huang, Chi-Shiang Chiou, Shih-Wei Liu, Tsun-Yen Wu Citations: 0 2.5D CoWoS advanced packaging has been broadly used in applications of artificial intelligence (AI) and high-performanc
Query: large language model inference gpu datacenter Authors: Giacomo Brunetta, V. Sastry, Xingfu Wu, Valerie Taylor, Michael E. Papka Citations: 0 Serving large language models (LLMs) is highly resource-intensive and requires specialized hardware acceleration. While GPUs remain
Query: mixture of experts inference serving Authors: Can Hankendi, Rana Shahout, Minlan Yu, A. Coskun Citations: 1 Large language model (LLM) inference has become a dominant workload in modern data centers, driving significant GPU utilization and energy consumption. While prior s
Query: large language model inference gpu datacenter Authors: Zhiye Song, Kyungmi Lee, E. K. Lee, Xin Zhang, Tamar Eilam Citations: 0 We present EnergyLens, an end-to-end framework for energy-aware large language model (LLM) inference optimization. As LLMs scale, predicting and r
Query: mixture of experts inference serving Authors: Ananya Hegde, Akshata Kumble, Ravi Gupta Citations: 0 Mixture-of-Experts (MoE) architectures reduce inference cost by activating only a sparse subset of parameters per token. However, when these models exceed single-GPU memory,
- 2026-04-25Research paper: Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
Query: mixture of experts inference serving Authors: A. Bambhaniya, Geonhwa Jeong, Jason Park, Jiecao Yu, Jaewon Lee Citations: 0 Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportion
Query: photonic computing transformer accelerator Authors: S. Afifi, O. Alo, I. Thakkar, S. Pasricha Citations: 0 Transformers achieve state-of-the-art performance in natural language processing, vision, and scientific computing, but demand high computation and memory. To address
- 2026-04-08Research paper: Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
Query: large language model inference gpu datacenter Authors: Mohammad Siavashi, Mariano Scazzariello, Gerald Q. Maguire, Dejan Kostic, Marco Chiesa Citations: 0 Large Language Model (LLM) inference is rapidly becoming a core datacenter service, yet current serving stacks keep th
- 2026-03-06Research paper: A Bandwidth-Efficient Bit-Serial Systolic Accelerator with Bit-Matrix Based Dataflow
Query: high bandwidth memory artificial intelligence accelerator Authors: Jianghui Wu, Weixing Li, Siyi Deng, Wujiao He, Yonghao Li Citations: 0
Query: high bandwidth memory artificial intelligence accelerator Authors: Ramakrishna Penaganti Citations: 0 Transformer models have become the most popular type of architecture in deep learning today. They can do the best work in natural language processing, computer vision, and
Query: high bandwidth memory artificial intelligence accelerator Authors: Riccardo Tedeschi, Luigi Ghionda, A. Nadalini, Yvan Tortorella, A. Prasad Citations: 0 Low Earth Orbit (LEO) constellations are revolutionizing the space sector, with on-board Artificial Intelligence (AI) b
Query: high bandwidth memory artificial intelligence accelerator Authors: Zhan Li, Yuxian Jiang, Zhihan Zhang, Jiahui Huang, Qunkang Meng Citations: 2
Query: ai datacenter power cooling efficiency Authors: Nardos Belay Abera, Yize Chen Citations: 0 The AI datacenters are currently being deployed on a large scale to support the training and deployment of power-intensive large-language models (LLMs). Extensive amount of computati
Query: mixture of experts inference serving Authors: Adrian Zhao, Zhenkun Cai, Zhenyu Song, Lin Yu, Haozheng Fan Citations: 1 Mixture-of-Experts (MoE) has recently emerged as the mainstream architecture for efficiently scaling large language models while maintaining near-constant
Query: photonic computing transformer accelerator Authors: Xiaofeng Zou, Limingrui Wan, Ziqian Zeng, Huiping Zhuang, Mingkui Tan Citations: 0
- 2025-12-31Research paper: Waste-to-Energy-Coupled AI Data Centers: Cooling Efficiency and Grid Resilience
Query: ai datacenter power cooling efficiency Authors: Qi He, Chunyu Qu Citations: 21 AI data center expansion is increasingly constrained by grid interconnection capacity, driven in part by electricity-intensive cooling demand. This paper evaluates a coupled Waste-to-Energy (WtE
Query: large language model inference gpu datacenter Authors: VG Jithin, PS Ditto Citations: 0 The proliferation of GPU-accelerated workloads, particularly in artificial intelligence and large language model (LLM) inference, has created unprecedented demand for efficient GPU reso
Query: mixture of experts inference serving Authors: Kexin Chu, Dawei Xiang, Zixu Shen, Yiwei Yang, Zechen Liu Citations: 3 Mixture-of-Experts (MoE) has become a practical architecture for scaling LLM capacity while keeping per-token compute modest, but deploying MoE models on a
Query: ai datacenter power cooling efficiency Authors: Pavana Prakash, R. P. Hong Enriquez, S. Serebryakov, David Grant, Wesley Brewer Citations: 3 Modern datacenters operate at unprecedented scale, supporting HPC and AI workloads while consuming hundreds of megawatts of power. T
Query: ai datacenter power cooling efficiency Authors: A. Lakshminarayana, Lochan Sai Reddy Chinthaparthy, Anto Barigala, R. Suthar, D. Agonafer Citations: 0 Datacenter industry is advancing rapidly over the years, cooling technologies have been advancing with increasing computin
Query: ai datacenter power cooling efficiency Authors: Braxton J. Smith, S. Pundla, D. Agonafer Citations: 0 The increased demand for computational power in modern computing components, coupled with the advancements in packaging technologies used to achieve this improvement, has
Query: photonic computing transformer accelerator Authors: Hanqing Zhu, Zhican Zhou, Shupeng Ning, Xuhao Wu, Ray T. Chen Citations: 0 Photonic computing has emerged as a promising substrate for accelerating the dense linear-algebra operations at the heart of AI, but its adoption
Query: ai datacenter power cooling efficiency Authors: Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Ricardo Bianchini Citations: 6 The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs.
- 2025-09-26Research paper: Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Query: mixture of experts inference serving Authors: Naibin Gu, Zhenyu Zhang, Yuchen Feng, Yilong Chen, Peng Fu Citations: 8 Mixture-of-Experts (MoE) models typically fix the number of activated experts $k$ at both training and inference. However, real-world deployments often fac
Query: mixture of experts inference serving Authors: Rui Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo Citations: 27 Mixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational complexi
Query: high bandwidth memory artificial intelligence accelerator Authors: Feng Jiao, Siying Bi, Shu Wang, Zhong Ma, Jin Huang Citations: 0 With the rapid advancement of intelligent technologies and their widespread deployment in edge scenarios, energy efficiency has become a crit
Query: photonic computing transformer accelerator Authors: Yi Li, Zijian Ye, Xiangqu Fu, Song-jian Wang, Shucheng Du Citations: 1 Vision Transformers (ViTs) are new foundation models for vision applications. Edge-deploying ViTs to realize energy-saving, low-latency, and high-perf
Query: ai datacenter power cooling efficiency Authors: Yang Liu, Rishav Roy, David J. Apigo, Manohar Bongarala, Syed Faisal Citations: 1 The increasing thermal demands of high-performance computing in AI datacenters, 5G radio access networks (RAN), and edge compute nodes necessit
Query: high bandwidth memory artificial intelligence accelerator Authors: Shunsuke Tonouchi, Yuji Sakurazawa, Ayaka Takeguchi, Keiichi Kasuga, Yusuke Sera Citations: 2 Artificial Intelligence (AI) technology in natural language processing and synthetic image generation have signi
Query: ai datacenter power cooling efficiency Authors: Pratikkumar Dilipkumar, Pratikkumar Dilipkumar Patel Citations: 6 Artificial Intelligence is fundamentally transforming datacenter infrastructure management, creating unprecedented opportunities for performance optimization,
Query: mixture of experts inference serving Authors: Shaoyu Wang, Guangrong He, Geon-Woo Kim, Yanqi Zhou, S. Park Citations: 3 Mixture-of-Experts (MoE) architectures offer the promise of larger model capacity without the prohibitive costs of fully dense designs. However, in real-
- 2025-05-01Research paper: AMD Instinct MI300X: A Generative AI Accelerator and Platform Architecture
Query: high bandwidth memory artificial intelligence accelerator Authors: Alan Smith, V. Alla Citations: 0 AMD Instinct MI300X sets a new benchmark in generative artificial intelligence (AI) acceleration, combining architectural innovation with advanced system integration to tack
Query: high bandwidth memory artificial intelligence accelerator Authors: Yuankang Zhao, Arne Heittman, Emre Neftci Citations: 2 Recent development in neuromorphic hardware focused on exploiting the temporal and spatial sparsity of Spiking Neural Networks (SNNs). Event driven acc
- 2025-04-24Research paper: Energy Considerations of Large Language Model Inference and Efficiency Optimizations
Query: large language model inference gpu datacenter Authors: Jared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk, Sasha Luccioni Citations: 75 As large language models (LLMs) scale in size and adoption, their computational and environmental costs continue to rise. Prior ben
Query: mixture of experts inference serving Authors: Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo Citations: 35 Mixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational comp
Query: photonic computing transformer accelerator Authors: Shiyue Hua, Erwan Divita, Shanshan Yu, Bo Peng, C. Roques-Carmes Citations: 157 Integrated photonics, particularly silicon photonics, have emerged as cutting-edge technology driven by promising applications such as short-
- 2025-03-31Research paper: HyAtten: Hybrid Photonic-Digital Architecture for Accelerating Attention Mechanism
Query: photonic computing transformer accelerator Authors: Huize Li, Dan Chen, Tulika Mitra Citations: 0 The wide adoption and substantial computational resource requirements of attention-based Transformers have spurred the demand for efficient hardware accelerators. Unlike digit
Query: large language model inference gpu datacenter Authors: Ruibo Fan, Xiangrui Yu, Peijie Dong, Zeyu Li, G. Gong Citations: 33 Large Language Models (LLMs) have demonstrated remarkable capabilities, but their immense scale poses significant challenges in terms of both memory a
Query: photonic computing transformer accelerator Authors: Bo Chen, T. Chang Citations: 3 This paper introduces the first low-power hardware accelerator for Spiking Transformers, an emerging alternative to traditional artificial neural networks. By modifying the base Spikformer m
Query: high bandwidth memory artificial intelligence accelerator Authors: C. Pappas, A. Prapas, T. Moschos, M. Kirtas, O. Asimopoulos Citations: 4 The ever-increasing volume of data demarcating from the exponential scale of Artificial Intelligence (AI) and Deep Learning (DL) mode
Query: photonic computing transformer accelerator Authors: Pingcheng Dong, Yonghao Tan, Xuejiao Liu, Peng Luo, Yu Liu Citations: 17 Recently, hybrid models integrating a CNN and a Transformer (ConvFormer), shown in Fig. 23.2.1, have achieved significant advancements in semantic s
Query: large language model inference gpu datacenter Authors: Yufeng Gu, Alireza Khadem, S. Umesh, Ning Liang, Xavier Servot Citations: 60 Large Language Model (LLM) inference uses an autoregressive manner to generate one token at a time, which exhibits notably lower operational
Query: photonic computing transformer accelerator Authors: Jiaqi Liu, Yiwen Ma Citations: 1 The demand for extensive computing resources and energy to support the increasing size of machine learning models has created a disparity between AI applications and the underlying hardwar
Query: photonic computing transformer accelerator Authors: Huize Li, Dan Chen, Tulika Mitra Citations: 0 The wide adoption and substantial computational resource requirements of attention-based Transformers have spurred the demand for efficient hardware accelerators. Unlike digit