FlashML-org/FreeToken
Local MoE model serving engine
FreeToken is an edge-native serving engine that enables running massive Mixture-of-Experts models locally on consumer hardware. It optimizes heterogeneous resources like GPUs and CPUs to deliver fast, interactive inference for frontier-scale open-weight models.
- Stars
- 13,700
- Contributors
- 8
- Forks
- 1,400
- Days trending in the last 30
- 2 days
- Mark
- Watch — 1,712 stars per contributor. An automatic mark on public metrics: above 500 is watch, above 2,000 possible hype (not a judgement of quality).