| 1 |
vllm |
92911 |
22794 |
Python |
2546 |
A high-throughput and memory-efficient inference and serving engine for LLMs |
2026-09-29T09:45:03Z |
| 2 |
LlamaFactory |
75177 |
9201 |
Python |
999 |
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024) |
2026-09-28T09:10:59Z |
| 3 |
colibri |
38234 |
4172 |
C |
54 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦 |
2026-09-28T19:47:37Z |
| 4 |
sglang |
36565 |
9204 |
Python |
901 |
SGLang is a high-performance serving framework for large language models and multimodal models. |
2026-09-29T09:44:30Z |
| 5 |
ms-swift |
15751 |
1706 |
Python |
463 |
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, …) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM5.3, Gemma4, Llava, Phi4, …) (AAAI 2025). |
2026-09-29T02:23:44Z |
| 6 |
TensorRT-LLM |
14737 |
2781 |
Python |
606 |
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. |
2026-09-29T09:38:39Z |
| 7 |
FreeToken |
13961 |
1380 |
Python |
204 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently. |
2026-09-26T11:09:07Z |
| 8 |
kimi-k3-in-c |
8788 |
1421 |
C |
4 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU. |
2026-09-22T14:29:39Z |
| 9 |
flashinfer |
6520 |
1510 |
Cuda |
353 |
FlashInfer: Kernel Library for LLM Serving |
2026-09-29T09:18:38Z |
| 10 |
MoeKoeMusic |
6378 |
402 |
Vue |
8 |
一款开源简洁高颜值的酷狗第三方客户端 An open-source, concise, and aesthetically pleasing third-party client for KuGou that supports Windows / macOS / Linux / Web :electron: |
2026-09-22T12:27:21Z |
| 11 |
Bangumi |
5997 |
167 |
TypeScript |
27 |
:electron: An unofficial https://bgm.tv ui first app client for Android and iOS, built with React Native. 一个无广告、以爱好为驱动、不以盈利为目的、专门做 ACG 的类似豆瓣的追番记录,bgm.tv 第三方客户端。为移动端重新设计,内置大量加强的网页端难以实现的功能,且提供了相当的自定义选项。 目前已适配 iOS / Android。 |
2026-09-29T08:46:08Z |
| 12 |
xtuner |
5204 |
451 |
Python |
243 |
A Next-Generation Training Engine Built for Ultra-Large MoE Models |
2026-09-27T15:59:41Z |
| 13 |
fastllm |
5083 |
500 |
C++ |
318 |
fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tps,多并发可达60+。 |
2026-09-25T14:14:41Z |
| 14 |
trace.moe |
5034 |
261 |
None |
0 |
Timestamp Retrieval for Anime Clips Everywhere |
2026-09-23T19:33:47Z |
| 15 |
GLM-4.5 |
4426 |
477 |
Python |
27 |
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models |
2026-02-01T08:28:10Z |
| 16 |
flash-moe |
4154 |
511 |
Objective-C |
10 |
Running a big model on a small laptop |
2026-03-19T17:21:57Z |
| 17 |
Moeditor |
4100 |
263 |
JavaScript |
106 |
(discontinued) Your all-purpose markdown editor. |
2020-07-07T01:08:32Z |
| 18 |
Moe-Counter |
3096 |
301 |
JavaScript |
3 |
Moe counter badge with multiple themes! - 多种风格可选的萌萌计数器 |
2026-04-16T03:39:37Z |
| 19 |
moemail |
2820 |
2593 |
TypeScript |
50 |
A cute temporary email service built with NextJS + Cloudflare technology stack 🎉 | 一个基于 NextJS + Cloudflare 技术栈构建的可爱临时邮箱服务🎉 |
2026-08-29T10:20:42Z |
| 20 |
MoeGoe |
2424 |
240 |
Python |
28 |
Executable file for VITS inference |
2023-08-22T07:17:37Z |
| 21 |
MoE-LLaVA |
2322 |
139 |
Python |
65 |
【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models |
2025-07-15T07:59:33Z |
| 22 |
MoBA |
2190 |
161 |
Python |
13 |
MoBA: Mixture of Block Attention for Long-Context LLMs |
2025-04-03T07:28:06Z |
| 23 |
ICEdit |
2103 |
109 |
Python |
24 |
[NeurIPS 2025] Image editing is worth a single LoRA! 0.1% training data for fantastic image editing! Surpasses GPT-4o in ID persistence~ MoE ckpt released! Only 4GB VRAM is enough to run! |
2025-12-19T19:08:02Z |
| 24 |
DeepSeek-MoE |
1978 |
311 |
Python |
18 |
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models |
2024-01-16T12:18:10Z |
| 25 |
fastmoe |
1860 |
206 |
Python |
25 |
A fast MoE impl for PyTorch |
2025-02-10T06:04:33Z |
| 26 |
OpenMoE |
1701 |
85 |
Python |
6 |
A family of open-sourced Mixture-of-Experts (MoE) Large Language Models |
2024-03-08T15:08:26Z |
| 27 |
paimon-moe |
1533 |
282 |
JavaScript |
316 |
Your best Genshin Impact companion! Help you plan what to farm with ascension calculator and database. Also track your progress with todo and wish counter. |
2026-09-23T08:12:09Z |
| 28 |
uccl |
1533 |
177 |
C++ |
57 |
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven) |
2026-09-28T05:59:42Z |
| 29 |
diy-llm |
1431 |
153 |
Jupyter Notebook |
0 |
Covers pre-training data, Tokenizer, Transformer, MoE,distributed training, Scaling Laws, inference & alignment .6 progressive code assignments for full-stack LLM learning | 涵盖预训练数据、分词器、Transformer、MoE、分布式训练、缩放定律、推理与对齐,6 项渐进代码作业,掌握 LLM 全栈知识 |
2026-09-10T13:03:41Z |
| 30 |
SpikingBrain-7B |
1383 |
191 |
Python |
10 |
Spiking Brain-inspired Large Models, integrating hybrid efficient attention, MoE modules and spike encoding into its architecture |
2026-05-14T09:52:16Z |
| 31 |
moepush |
1368 |
434 |
TypeScript |
15 |
一个基于 NextJS + Cloudflare 技术栈构建的可爱消息推送服务, 支持多种消息推送渠道✨ |
2025-05-10T11:42:44Z |
| 32 |
SmartImage |
1349 |
82 |
C# |
4 |
Reverse image search tool (SauceNao, IQDB, Ascii2D, trace.moe, and more) |
2026-09-06T01:55:53Z |
| 33 |
MOE |
1321 |
138 |
C++ |
170 |
A global, black box optimization engine for real world metric optimization. |
2023-03-24T11:00:32Z |
| 34 |
Strata |
1284 |
154 |
C++ |
19 |
Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input. |
2026-09-29T09:41:22Z |
| 35 |
mixture-of-experts |
1253 |
112 |
Python |
6 |
PyTorch Re-Implementation of “The Sparsely-Gated Mixture-of-Experts Layer” by Noam Shazeer et al. https://arxiv.org/abs/1701.06538 |
2024-04-19T08:22:39Z |
| 36 |
MiniMind-in-Depth |
1209 |
91 |
None |
6 |
轻量级大语言模型MiniMind的源码解读,包含tokenizer、RoPE、MoE、KV Cache、pretraining、SFT、LoRA、DPO等完整流程 |
2025-06-16T14:13:15Z |
| 37 |
MoeMemosAndroid |
1195 |
138 |
Kotlin |
97 |
An app to help you capture thoughts and ideas |
2026-09-28T07:05:40Z |
| 38 |
Uni-MoE |
1118 |
72 |
Python |
27 |
Uni-MoE: Lychee’s Large Multimodal Model Family. |
2026-08-06T10:50:45Z |
| 39 |
Aria |
1088 |
92 |
Jupyter Notebook |
33 |
Codebase for Aria - an Open Multimodal Native MoE |
2025-01-22T03:25:37Z |
| 40 |
Tutel |
1022 |
111 |
C |
55 |
Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4 |
2026-09-15T05:55:02Z |
| 41 |
llama-moe |
1003 |
60 |
Python |
6 |
⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024) |
2024-12-06T04:47:07Z |
| 42 |
Time-MoE |
1002 |
116 |
Python |
14 |
[ICLR 2025 Spotlight] Official implementation of “Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts” |
2026-03-21T16:00:55Z |
| 43 |
MoeTTS |
987 |
72 |
None |
0 |
Speech synthesis model /inference GUI repo for galgame characters based on Tacotron2, Hifigan, VITS and Diff-svc |
2023-03-03T07:30:05Z |
| 44 |
moebius |
976 |
54 |
JavaScript |
40 |
Modern ANSI & ASCII Art Editor |
2024-05-02T15:54:35Z |
| 45 |
cudnn-frontend |
953 |
299 |
Python |
107 |
cuDNN Frontend is NVIDIA’s modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels. |
2026-09-29T05:37:59Z |
| 46 |
ComfyUI-QwenVL |
898 |
145 |
Python |
2 |
ComfyUI-QwenVL custom node: Integrates the Qwen-VL series, including Qwen2.5-VL, Qwen3-VL, Qwen3.5-VL, Qwen3.6-VL (MoE), and Qwen3.8-VL, with GGUF support for advanced multimodal AI in text generation, image understanding, and video analysis. |
2026-09-14T07:39:09Z |
| 47 |
MoeMemos |
891 |
80 |
Swift |
81 |
An app to help you capture thoughts and ideas |
2026-09-26T03:43:14Z |
| 48 |
MoePeek |
835 |
58 |
Swift |
8 |
A lightweight macOS selection translator built with pure Swift 6, featuring on-device Apple Translate for privacy, only 5MB install size and stable ~50MB memory usage. 一款轻量级 macOS 划词翻译工具,纯 Swift 6 开发,设备端 Apple 翻译保护隐私,安装体积仅 5MB,后台运行内存稳定约 50MB |
2026-09-28T07:07:33Z |
| 49 |
Adan |
822 |
71 |
Python |
6 |
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models |
2025-06-08T14:35:41Z |
| 50 |
Hunyuan-A13B |
820 |
118 |
Python |
35 |
Tencent Hunyuan A13B (short as Hunyuan-A13B), an innovative and open-source LLM built on a fine-grained MoE architecture. |
2025-07-08T08:45:27Z |
| 51 |
DeepSeek-671B-SFT-Guide |
815 |
97 |
Python |
1 |
An open-source solution for full parameter fine-tuning of DeepSeek-V3/R1 671B, including complete code and scripts from training to inference, as well as some practical experiences and conclusions. (DeepSeek-V3/R1 满血版 671B 全参数微调的开源解决方案,包含从训练到推理的完整代码和脚本,以及实践中积累一些经验和结论。) |
2025-03-13T03:51:33Z |
| 52 |
halo |
779 |
37 |
Python |
6 |
Halo is an open-source framework built by White Circle for training large language and multimodal models |
2026-09-28T22:00:54Z |
| 53 |
SwiftLM |
775 |
54 |
Swift |
3 |
⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app. |
2026-09-28T15:42:14Z |
| 54 |
moe-theme.el |
773 |
68 |
Emacs Lisp |
15 |
A customizable colorful eye-candy theme for Emacser. Moe, moe, kyun! |
2026-08-19T15:54:08Z |
| 55 |
sonic-moe |
772 |
107 |
Python |
8 |
Accelerating MoE with IO and Tile-aware Optimizations |
2026-08-29T09:26:22Z |
| 56 |
MixtralKit |
770 |
76 |
Python |
11 |
A toolkit for inference and evaluation of ‘mixtral-8x7b-32kseqlen’ from Mistral AI |
2023-12-15T19:10:55Z |
| 57 |
YOLO-Master |
740 |
158 |
Python |
29 |
[CVPR2026]🚀🚀🚀Official code for the paper “YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection.” (YOLO = You Only Look Once) 🔥🔥🔥 |
2026-09-28T16:42:41Z |
| 58 |
moe |
728 |
36 |
Nim |
42 |
A command line based editor inspired by Vim. Written in Nim. |
2026-09-28T19:55:36Z |
| 59 |
pegainfer |
712 |
110 |
Rust |
51 |
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2 |
2026-09-28T23:14:44Z |
| 60 |
Awesome-Mixture-of-Experts-Papers |
672 |
49 |
None |
3 |
A curated reading list of research in Mixture-of-Experts(MoE). |
2024-10-30T07:48:14Z |
| 61 |
MoeList |
660 |
23 |
Kotlin |
26 |
Another unofficial Android MAL client |
2026-08-09T08:38:47Z |
| 62 |
moedict-webkit |
653 |
98 |
Objective-C |
1 |
萌典網站 |
2026-08-14T03:36:14Z |
| 63 |
vtbs.moe |
639 |
36 |
Vue |
34 |
Virtual YouTubers in bilibili |
2025-07-31T13:39:09Z |
| 64 |
satania.moe |
619 |
55 |
HTML |
3 |
Satania IS the BEST waifu, no really, she is, if you don’t believe me, this website will convince you |
2022-10-09T23:19:01Z |
| 65 |
moebooru |
613 |
81 |
Ruby |
30 |
Moebooru, a fork of danbooru1 that has been heavily modified |
2026-09-22T23:38:32Z |
| 66 |
Chinese-Mixtral |
612 |
43 |
Python |
0 |
中文Mixtral混合专家大模型(Chinese Mixtral MoE LLMs) |
2026-04-19T00:59:54Z |
| 67 |
moebius |
607 |
42 |
Elixir |
3 |
A functional query tool for Elixir |
2024-10-23T18:55:45Z |
| 68 |
BigMoeOnEdge |
596 |
63 |
C++ |
23 |
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, lossless, on stock llama.cpp |
2026-09-29T07:37:55Z |
| 69 |
mixture-of-kittens |
595 |
78 |
Python |
4 |
Mixture-of-experts (MoE) training megakernel for NVL72s |
2026-08-14T03:44:18Z |
| 70 |
BitSoulStockSkill |
582 |
53 |
Python |
0 |
由BitSoul出品的A股市场全能Skill,自带免费历史数据,内置100+行业主流因子,完整的回测框架,基于MOE架构的股票筛选与买卖判断,更提供因子挖矿等趣味接口,欢迎安装试用,也欢共同开发交流! |
2026-03-21T08:19:00Z |
| 71 |
MoeGoe_GUI |
568 |
69 |
C# |
8 |
GUI for MoeGoe |
2023-08-22T07:32:08Z |
| 72 |
trace.moe-telegram-bot |
560 |
78 |
TypeScript |
0 |
This Telegram Bot can tell the anime when you send an screenshot to it |
2026-09-23T12:42:47Z |
| 73 |
ARIS-in-AI-Offer |
555 |
20 |
Python |
3 |
Bilingual (中文+EN) ML / LLM / diffusion / agent interview cheat sheets for AI 秋招 — generated by ARIS /interview-cheatsheet, rendered by /render-html into single-file HTML, reads anywhere — plus a CV→DBLP-fact-checked academic homepage generator and hand-authored long-form blogs 🌱 |
2026-09-28T07:48:55Z |
| 74 |
moerail |
547 |
43 |
JavaScript |
22 |
铁路车站代码查询 × 动车组交路查询 |
2025-08-13T12:55:25Z |
| 75 |
kaggle-tpu-lab |
546 |
72 |
Python |
10 |
Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi. |
2026-09-15T14:33:29Z |
| 76 |
vLLM-Moet |
541 |
49 |
Sass |
11 |
A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 “delta” cache that recovers precision — matching the official (NV)FP4 checkpoint’s quality on consumer Blackwell cards |
2026-09-25T21:05:30Z |
| 77 |
Moebius |
540 |
45 |
Python |
3 |
[ECCV 2026] Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance |
2026-08-12T03:16:36Z |
| 78 |
LPLB |
534 |
43 |
Python |
1 |
An early research stage expert-parallel load balancer for MoE models based on linear programming. |
2025-11-19T07:20:35Z |
| 79 |
FreeMoe |
485 |
8 |
None |
16 |
Unlock App Vip |
2026-01-30T05:50:54Z |
| 80 |
step_into_llm |
480 |
127 |
Jupyter Notebook |
27 |
MindSpore online courses: Step into LLM |
2025-12-22T11:46:46Z |
| 81 |
apex-quant |
475 |
34 |
Shell |
11 |
Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization |
2026-08-17T09:15:33Z |
| 82 |
Lvllm |
462 |
41 |
Python |
0 |
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models. |
2026-09-22T03:49:46Z |
| 83 |
MoeSR |
457 |
13 |
JavaScript |
7 |
An application specialized in image super-resolution for ACGN illustrations and Visual Novel CG. 专注于插画/Galgame CG等ACGN领域的图像超分辨率的应用 |
2026-08-10T15:16:16Z |
| 84 |
zeraix |
445 |
13 |
TypeScript |
0 |
Open-source local AI workspace — advancing on-device inference. |
2026-09-24T08:30:35Z |
| 85 |
DiT-MoE |
439 |
20 |
Python |
7 |
Scaling Diffusion Transformers with Mixture of Experts |
2024-09-09T02:12:12Z |
| 86 |
moe-sticker-bot |
431 |
75 |
Go |
39 |
A Telegram bot that imports LINE/kakao stickers or creates/manages new sticker set. |
2024-06-06T15:28:28Z |
| 87 |
awesome-moe-inference |
427 |
19 |
None |
0 |
Curated collection of papers in MoE model inference |
2026-03-12T01:59:19Z |
| 88 |
MOE |
422 |
80 |
Java |
18 |
Make Opensource Easy - tools for synchronizing repositories |
2022-06-20T22:41:08Z |
| 89 |
hydra-moe |
415 |
16 |
Python |
10 |
None |
2023-11-02T22:53:15Z |
| 90 |
MoeLoaderP |
413 |
28 |
C# |
12 |
🖼二次元图片下载器 Pics downloader for booru sites,Pixiv.net,Bilibili.com,Konachan.com,Yande.re , behoimi.org, safebooru, danbooru,Gelbooru,SankakuComplex,Kawainyan,MiniTokyo,e-shuushuu,Zerochan,WorldCosplay ,Yuriimg etc. |
2025-05-19T13:20:58Z |
| 91 |
ds4-on-spark |
407 |
29 |
Shell |
7 |
Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support |
2026-08-27T04:25:21Z |
| 92 |
Awesome-Efficient-Arch |
407 |
33 |
None |
0 |
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models |
2025-11-11T09:47:37Z |
| 93 |
slotstream |
406 |
27 |
Swift |
1 |
Run a 105 GB AI model on a Mac that can’t hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients. |
2026-09-28T22:11:28Z |
| 94 |
nmoe |
399 |
33 |
Python |
2 |
MoE training for Me and You and maybe other people |
2026-03-15T22:23:47Z |
| 95 |
st-moe-pytorch |
388 |
35 |
Python |
4 |
Implementation of ST-Moe, the latest incarnation of MoE after years of research at Brain, in Pytorch |
2024-06-17T00:48:47Z |
| 96 |
WThermostatBeca |
373 |
70 |
C++ |
4 |
Open Source firmware replacement for Tuya Wifi Thermostate from Beca and Moes with Home Assistant Autodiscovery |
2023-08-26T22:10:38Z |
| 97 |
pixiv.moe |
371 |
41 |
TypeScript |
0 |
😘 A pinterest-style layout site, shows illusts on pixiv.net order by popularity. |
2023-03-08T06:54:34Z |
| 98 |
MoE-Infinity |
366 |
41 |
Python |
5 |
PyTorch library for cost-effective, fast and easy serving of MoE models. |
2026-09-29T09:41:22Z |
| 99 |
MOEAFramework |
363 |
128 |
Java |
0 |
A Free and Open Source Java Framework for Multiobjective Optimization |
2026-01-21T16:26:02Z |
| 100 |
pith-train |
354 |
38 |
Python |
4 |
Compact and Agent-Native MoE Training System |
2026-09-23T19:10:03Z |