| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuseDistributed & Parallel
Narrow Precision Training
Modeling Front
Model Optimization for Efficient Inference
Post-Training Statistical Calibration for Higher Activation Sparsity, [ENLSP 2024 Spotlight 7, Paper, Oral, Code, Integrated]
Pre-LLM explosion — Unified HuggingFace Trainer for Joint Pruning, Quantization, and Distillation (JPQD), integrating OpenVINO NNCF and runtime. 16× more BERT serving throughput on Xeon Sapphire Rapids. See MLPerf Inference 3.0 submission. Applicable to vision, audio models.
Perhaps useful: dlbp, dockerhub, HuggingFace
mxfp8/nvfp4 training - from concept to implementation (cuBLASLt + Microxcaling).
Python 5
Hands-on Megatron-LM tutorials on ablating parallelism and scaling trends. DP → ZeRO → TP → SP → CP → PP → VPP → EP
Shell 3
Plots & Takeways from MLPerf Training v6.0 on new MoE workloads (DeepSeek-v3, GPT-OSS), scaling efficiency and MXFP4 recipe debut.
Python
Implementation of MoE & Expert-Parallel (EP) communication using Pytorch Symmetric Memory.
Python
Hands-on MoE ablations (load balancing, resolution, shared experts, LatentMoE, etc.) using a subclassed HF Transformer on a single GPU.
Python
| Back | FazBrowse Home | New Git URL |