jundot/omlx
macOS LLM inference server
A local large language model inference server optimized for Apple Silicon that features continuous batching and tiered caching. It is managed through a macOS menu bar application to allow users to control model memory and context limits.
- Stars
- 22,100
- Stars gained in 7 days
- +300
- Contributors
- 269
- Forks
- 1,900
- Days trending in the last 30
- 1 days