FastLM
We develop efficient models and systems for large-scale, distributed, parallel, sparsity senarios.
Popular repositories Loading
-
CXL-SpecKV
CXL-SpecKV Public[FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
-
CSV-Decode
CSV-Decode PublicCSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
Python 13
-
tinyserve-vllm
tinyserve-vllm Public[ACM MM 2025 Oral] TinyServe: Query-Aware Page Allocation Optimization
Repositories
Showing 10 of 12 repositories
- CXL-SpecKV Public
[FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
- CSV-Decode Public
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
-
Top languages
Loading…
Most used topics
Loading…