Date & Time:
October 15, 2024 12:30 pm – 1:30 pm
Location:
JCL 298
10/15/2024 12:30 PM 10/15/2024 01:30 PM America/Chicago Kuntai DU (UChicago)- Optimizing Communication for Distributed LLM Inference JCL 298

Abstract: Previous work has identified GPU memory capacity as the primary bottleneck in LLM inference, necessitating effective KV cache management strategies. However, based on a long-term collaboration with the most popular open-source serving engine vLLM, we observe that the landscape is shifting toward distributed LLM inference due to new emerging trends such as long-context KV cache reuse, disaggregated prefilling, and multi-modal LLM inference. The central challenge thus evolves to efficient KV cache communication mechanisms. This talk explores potential solutions for optimizing communication in distributed systems and argues that effective communication requires two new roles: an orchestrator and a KV cache-store. These roles must collaborate closely to meet stringent service-level objectives, paving the way for scalable and efficient distributed LLM-serving systems

Speakers

Kuntai Du

PhD Candidate

Kuntai Du is a 6th-year PhD from UChicago. His research focus is data transfer for distributed inference systems, including analytic-aware video streaming in distributed video analytic settings and effective KV cache transfer for distributed LLM inference. He is the recipient of the Siebel Scholarship.

Related News & Events

UChicago CS News

UChicago CS Researchers Shine at UIST 2024 with Papers, Posters, Workshops and Demonstrations

Oct 10, 2024
UChicago CS News

UChicago Scientists Receive Grant to Expand Global Data Management Platform, Globus

Oct 03, 2024
UChicago CS News

UChicago Researchers Demonstrate the Quantifiable Uniqueness of Former President Donald Trump’s Language Use

Sep 30, 2024
UChicago CS News

Five UChicago CS students named to Siebel Scholars class of 2025

Sep 20, 2024
UChicago CS News

NSF and Simons Foundation launch $20 million National AI Research Institute in Astronomy

Sep 18, 2024
In the News

Data Ecology: A Socio-Technical Approach to Controlling Dataflows

Sep 18, 2024
UChicago CS News

Ph.D. Student Shawn Shan Named MIT Technology Review’s 35 Innovators Under 35 and Innovator of the Year

Sep 16, 2024
UChicago CS News

Ben Zhao Named to TIME Magazine’s TIME100 AI List

Sep 05, 2024
UChicago CS News

Ian Foster and Rick Stevens Named to HPCwire’s 35 Legends List

Aug 28, 2024
UChicago CS News

University of Chicago to Develop Software for Effort to Create a National Quantum Virtual Laboratory

Aug 28, 2024
UChicago CS News

New Classical Algorithm Enhances Understanding of Quantum Computing’s Future

Aug 27, 2024
UChicago CS News

Decoding Content Moderation: Analyzing Policy Variations Across Top Online Platforms

Aug 26, 2024
arrow-down-largearrow-left-largearrow-right-large-greyarrow-right-large-yellowarrow-right-largearrow-right-smallbutton-arrowclosedocumentfacebookfacet-arrow-down-whitefacet-arrow-downPage 1CheckedCheckedicon-apple-t5backgroundLayer 1icon-google-t5icon-office365-t5icon-outlook-t5backgroundLayer 1icon-outlookcom-t5backgroundLayer 1icon-yahoo-t5backgroundLayer 1internal-yellowinternalintranetlinkedinlinkoutpauseplaypresentationsearch-bluesearchshareslider-arrow-nextslider-arrow-prevtwittervideoyoutube