Where the Time Goes: Analysis of a Public LLM Serving System
- Büsra Karatay Demiray,
- ,
- Benoît Garbinato,
- ,
- Pamela Delgado
- HES-SO Valais-Wallis,
- ,
- ,
- ,
- University of Lausanne,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 171-182 (12 pages)Publication milestones
- Published - 28/04/2026
Publication status
Published - 28/04/2026
Publisher
Association for Computing Machinery, United StatesISBN (Print)
979-8-4007-2605-7Publication IDs
- Scopus: 105038690056
Host publication title
Proceedings of the Sixth European Workshop on Machine Learning and Systems, EuroMLSys 2026, Edinburgh, Scotland, UK, April 27-30, 2026Abstract
In this study, we present a characterization of serving traces collected from Public Al's serving of Apertus, an open source Large Language Model (LLM). The trace spans roughly five months (September 2025-January 2026) and contains 337K requests. We analyzed request sizes, token and timing behaviour, latency, model-size effects, and temporal patterns. Our findings show insights that do not align with common assumptions; (1) time-to-first-token is often driven by queuing rather than prefill compute, especially for small requests; (2) the 8B and 70B models show nearly the same user-perceived latency despite a 9× parameter gap; (3) a substantial fraction of requests are prefill/queuing-dominated rather than decode-dominated; and; (4) observable input features are weak predictors of output, which makes size-aware scheduling difficult at arrival time. As a contribution to the research community, we will publish this anonymized trace along with its analysis.
Publication metrics
PlumX, opens in new tab
Captures
2
Funding Details
This work was supported by the Swiss National Science Foundation (SNSF) under Grant 10001932(DEEP: Deep Learning Resource-Efficient GPU Orchestrator),International Co-Investigator Scheme.
Access to documents
Related Event
Title
Computer Systems
Event type
ConferenceDegree of recognition
International eventDate
27/04/2026 - 30/04/2026Location
EdinburghUnited Kingdom
