semantic-scholar
Semantic Scholar1 events on 2026-04-08role: thematic30d fetch windowcadence: on demandaccess: keyless
Extracts: paper title, abstract snippet · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public endpoint; no project credential required.
Last ingested: 2026-08-12 · run status: on demand
Archive source — full history has value. Use pagination to browse older records.
- 2026-04-08Research paper: Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
Query: large language model inference gpu datacenter Authors: Mohammad Siavashi, Mariano Scazzariello, Gerald Q. Maguire, Dejan Kostic, Marco Chiesa Citations: 0 Large Language Model (LLM) inference is rapidly becoming a core datacenter service, yet current serving stacks keep th