NVIDIA/Model-Optimizer
Model optimization library
A unified library providing state-of-the-art techniques such as quantization, pruning, and distillation to compress deep learning models. It optimizes inference speed for deployment frameworks like TensorRT-LLM and vLLM.
- Stars
- 4,000
- Contributors
- 104
- Forks
- 626
- Days trending in the last 30
- 1 days