| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
🚀🚀 Efficient implementations of Native Sparse Attention
This is the repo for the paper Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining.
We introduce UltraLLaDA , a scaled variant of LLaDA-8B-Base that extends the context length up to 128K tokens with light-weight post-training, enabling long-context comprehension and generation.
Python 15
SGLang is a high-performance serving framework for large language models and multimodal models.
[ICML 2024, ICLR 2025, ICML 2025, ICML 2026] HexGen: LLM Serving over Heterogeneous GPUs; HexGen-2: PD Disaggregation over Heterogeneous GPUs; Demystifying Cost-Efficiency of LLM Serving over Heterogeneous GPUs; HexGen-3: Fully Disaggregated Serving and Resource Autoscaling over Heterogeneous GPUs.
A Text2SQL Serving System based on CHESS Framework
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
This organization has no public members. You must be a member to see who’s a part of this organization.
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |