deepseek-ai/DeepGEMM
GPU tensor core kernel library
DeepGEMM is a high-performance CUDA library providing optimized BLAS kernels and fused operations for large language model inference. It supports various data types and features like fused MoE and runtime compilation to accelerate AI workloads on GPUs.
- Language
- Cuda
- Stars
- 8,600
- Contributors
- 44
- Forks
- 1,400
- Days trending in the last 30
- 1 day
- Latest release
- v2.1.1.post3 · October 15, 2025
- Commits
- 197
- Last commit
- September 30, 2026
- Open issues
- 69
- Open pull requests
- 80