Skip to content
@FastLM

FastLM

We develop efficient models and systems for large-scale, distributed, parallel, sparsity senarios.

Popular repositories Loading

  1. CXL-SpecKV CXL-SpecKV Public

    [FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving

    C++ 38 10

  2. CSV-Decode CSV-Decode Public

    CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

    Python 13

  3. tinyserve-vllm tinyserve-vllm Public

    [ACM MM 2025 Oral] TinyServe: Query-Aware Page Allocation Optimization

    Shell 11 2

  4. SPI_VecDB SPI_VecDB Public

    [VecDB @ VLDB 2026] SPI: Query-Depth-Adaptive Indexing for Streaming RAG in Vector Databases

    Go 10

  5. HSGM HSGM Public

    [ICPADS 2025 Oral, *SEM 2025 Oral] HSGM: Hierarchical Segment-Graph Memory for Scalable Long-Text Semantics

    Python 8

  6. MKA MKA Public

    [ACM CF'26 Oral] MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning

    Python 8 1

Repositories

Showing 10 of 12 repositories
  • ForkServe Public
    FastLM/ForkServe's past year of commit activity
    Python 1 Apache-2.0 0 0 0 Updated Sep 13, 2026
  • KVLearn Public

    [SYSTOR 2026] To Keep or Not to Keep: Learning KV Cache Retention in Disaggregated LLM Serving Systems

    FastLM/KVLearn's past year of commit activity
    C++ 2 0 0 0 Updated Sep 4, 2026
  • CongShed Public

    [HPEC 2026] Breaking the Interconnect Bottleneck in LLM Serving via Congestion-Aware Scheduling over Heterogeneous Interconnects

    FastLM/CongShed's past year of commit activity
    C++ 2 0 0 0 Updated Aug 25, 2026
  • SPI_VecDB Public

    [VecDB @ VLDB 2026] SPI: Query-Depth-Adaptive Indexing for Streaming RAG in Vector Databases

    FastLM/SPI_VecDB's past year of commit activity
    Go 10 Apache-2.0 0 1 0 Updated Aug 17, 2026
  • CXL-SpecKV Public

    [FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving

    FastLM/CXL-SpecKV's past year of commit activity
    C++ 38 10 0 0 Updated Aug 7, 2026
  • OLIVE Public

    [UbiComp 2026] OLIVE: Online Low-Rank Incremental Learning for Efficient Adaptive Exoskeletons

    FastLM/OLIVE's past year of commit activity
    C++ 2 0 0 0 Updated Aug 7, 2026
  • CSV-Decode Public

    CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

    FastLM/CSV-Decode's past year of commit activity
    Python 13 0 0 0 Updated Aug 1, 2026
  • tinyserve-vllm Public

    [ACM MM 2025 Oral] TinyServe: Query-Aware Page Allocation Optimization

    FastLM/tinyserve-vllm's past year of commit activity
    Shell 11 2 0 0 Updated Jul 15, 2026
  • SemToken Public

    [*SEM 2026 Oral] SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Models

    FastLM/SemToken's past year of commit activity
    Python 5 0 0 0 Updated Jul 7, 2026
  • MKA Public

    [ACM CF'26 Oral] MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning

    FastLM/MKA's past year of commit activity
    Python 8 1 0 0 Updated Jun 29, 2026

Top languages

Loading…

Most used topics

Loading…