# vLLM Recipes > Per-model serving recipes for vLLM: hardware-tuned `vllm serve` commands, flag explanations, and known pitfalls. ## Core pages - [Homepage](https://recipes.vllm.ai): Recipe catalogue: pick a model, get a working `vllm serve` command. - [Browse all recipes](https://recipes.vllm.ai/browse): Flat alphabetical list of every recipe. - [JSON API](https://recipes.vllm.ai/models.json): Machine-readable list of every recipe and its metadata. ## Related properties - [vLLM](https://vllm.ai): project homepage and blog. - [vLLM documentation](https://docs.vllm.ai): full vLLM reference docs. - [vLLM on GitHub](https://github.com/vllm-project/vllm): source code and issues. ## Recipes ### MiMo (Xiaomi) - [XiaomiMiMo/MiMo-V2.6-Flash-RL](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.6-Flash-RL): Xiaomi's efficiency-balanced omnimodal MoE reasoning model (309B total / 15B active) with hybrid attention, FP8-compute/mxfp4-stored weights, 1M context, and a DFlash speculative decoder - [XiaomiMiMo/MiMo-V2.6-Pro-RL](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.6-Pro-RL): Xiaomi's flagship omnimodal MoE reasoning model (1.02T total / 42B active) with hybrid attention, FP8-compute/mxfp4-stored weights, 1M context, and a DFlash speculative decoder - [XiaomiMiMo/MiMo-V2.5-Pro](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.5-Pro): Xiaomi's flagship MoE reasoning model (1.02T total / 42B active) with hybrid attention, native FP8 weights, and Multi-Token Prediction - [XiaomiMiMo/MiMo-V2.5](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.5): MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture - [XiaomiMiMo/MiMo-V2-Flash](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2-Flash): Xiaomi's MoE reasoning model (309B total / 15B active) with hybrid attention and MTP for fast agentic workflows ### MiniMax - [MiniMaxAI/MiniMax-M3](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3): MiniMax M3 vision-language MoE (427B total / 26B active) for frontier coding, agent toolchains, and 1M-token reasoning via MSA sparse attention — native multimodal (image + video + computer use); BF16 plus MXFP8, NVIDIA Blackwell NVFP4, and AMD MI355X MXFP4 variants. Runs on NVIDIA (Hopper/Blackwell) and AMD CDNA4/CDNA3. - [MiniMaxAI/MiniMax-M2.5](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M2.5): MiniMax M2.5 MoE language model (230B total / 10B active) for coding, agent toolchains, and long-context reasoning — native FP8 checkpoint - [MiniMaxAI/MiniMax-M2.7](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M2.7): MiniMax M2.7 MoE language model (230B total / 10B active) — latest M2 release for coding, agent toolchains, and long-context reasoning with native FP8 - [MiniMaxAI/MiniMax-H3](https://recipes.vllm.ai/MiniMaxAI/MiniMax-H3): Open-weight general-purpose multimodal generation model — jointly generates 24 FPS video with native stereo audio from text, image, video, and audio references, served via vLLM-Omni - [MiniMaxAI/MiniMax-M2](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M2): MiniMax M2 MoE language model (230B total / 10B active) for coding, agent toolchains, and long-context reasoning — native FP8 checkpoint, with an NVFP4 variant for Blackwell - [MiniMaxAI/MiniMax-M2.1](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M2.1): MiniMax M2.1 MoE language model (230B total / 10B active) for coding, agent toolchains, and long-context reasoning — native FP8 checkpoint ### GLM (Z-AI) - [zai-org/GLM-5.1](https://recipes.vllm.ai/zai-org/GLM-5.1): GLM-5.1 refreshed version of GLM-5 — frontier-scale MoE language model (~744B total parameters) with MTP speculative decoding and thinking mode - [zai-org/GLM-5](https://recipes.vllm.ai/zai-org/GLM-5): GLM-5 frontier-scale MoE language model (~744B total parameters, 28.5T training tokens) with asynchronous RL infrastructure for reasoning, coding, and agentic tasks - [zai-org/GLM-5.2](https://recipes.vllm.ai/zai-org/GLM-5.2): GLM-5.2 — frontier-scale MoE language model (~743B total parameters, 39B active) with up to 5-token MTP speculative decoding and thinking mode - [zai-org/GLM-5.3](https://recipes.vllm.ai/zai-org/GLM-5.3): GLM-5.3 — Frontier Coding with Emergent Cyber Capabilities - [zai-org/GLM-4.7](https://recipes.vllm.ai/zai-org/GLM-4.7): GLM-4.7 MoE language model (~358B total parameters) with MTP speculative decoding, updated tool call parser, and reasoning support - [zai-org/GLM-5.3-Flash](https://recipes.vllm.ai/zai-org/GLM-5.3-Flash): GLM-5.3-Flash is a 320B-total / 18B-active multimodal MoE with hybrid KDA and sparse MLA attention, native FP8 weights, MTP, and a 1M-token context window. - [zai-org/GLM-4.5](https://recipes.vllm.ai/zai-org/GLM-4.5): GLM-4.5 MoE language model (~358B total parameters, BF16) with built-in MTP layers for speculative decoding and native tool calling - [zai-org/GLM-4.6](https://recipes.vllm.ai/zai-org/GLM-4.6): GLM-4.6 MoE language model (~357B total parameters, BF16) with MTP speculative decoding, native tool calling and reasoning - [zai-org/glm-4-9b-hf](https://recipes.vllm.ai/zai-org/glm-4-9b-hf): GLM-4 9B dense base language model with 8K context. - [zai-org/GLM-TTS](https://recipes.vllm.ai/zai-org/GLM-TTS): Z-AI's two-stage (AR + DiT flow-matching) zero-shot voice-cloning TTS for Chinese and English, served via vLLM-Omni through the OpenAI /v1/audio/speech API. Every request is conditioned on reference audio + its transcript. - [zai-org/GLM-GA](https://recipes.vllm.ai/zai-org/GLM-GA): GLM-GA dense vision-language model (~10B) — image and video understanding with 128K context and dedicated Glmga video processor (fps=2, up to 640 frames) - [zai-org/GLM-4.5V](https://recipes.vllm.ai/zai-org/GLM-4.5V): GLM-4.5 vision-language MoE model (~107B parameters, BF16) with image-text-to-text capability, 64K context, expert parallelism, and native FP8 - [zai-org/GLM-4.6V](https://recipes.vllm.ai/zai-org/GLM-4.6V): GLM-4.6 vision-language MoE model — image-text-to-text with 128K context, native FP8 checkpoint, and expert parallelism - [zai-org/GLM-ASR-Nano-2512](https://recipes.vllm.ai/zai-org/GLM-ASR-Nano-2512): Open-source speech recognition model (~2B) with strong dialect support (Cantonese and others) and robust low-volume speech transcription - [zai-org/GLM-Image](https://recipes.vllm.ai/zai-org/GLM-Image): Hybrid autoregressive + diffusion image generation model — text-to-image and image-to-image with strong text rendering and knowledge-intensive generation - [zai-org/GLM-OCR](https://recipes.vllm.ai/zai-org/GLM-OCR): GLM-OCR image-to-text model with built-in MTP speculative decoding for high-throughput OCR serving - [zai-org/Glyph](https://recipes.vllm.ai/zai-org/Glyph): Visual-text compression framework that renders long text into images and processes them with a reasoning VLM, scaling effective context length ### DeepSeek - [deepseek-ai/DeepSeek-V4.1-Flash](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4.1-Flash): DeepSeek V4.1 Flash vision-language MoE (552B backbone; 8B active per prompt token, 16B per output token) combining sliding-window plus compressed sparse attention with a two-level indexer, engram n-gram memory, hyper-connections, and a DSpark multi-token draft head. - [deepseek-ai/DeepSeek-R1](https://recipes.vllm.ai/deepseek-ai/DeepSeek-R1): DeepSeek-R1 is a 671B-parameter MoE reasoning model built on the DeepSeek-V3 architecture, trained with large-scale reinforcement learning for strong chain-of-thought capabilities. - [deepseek-ai/DeepSeek-V4-Pro](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Pro): DeepSeek V4 flagship MoE (1.6T total / 49B active) with hybrid CSA+HCA attention, manifold-constrained hyper-connections, Muon-trained on 32T+ tokens, and three-tier reasoning. - [deepseek-ai/DeepSeek-V4-Flash](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash): DeepSeek V4 MoE model with hybrid CSA+HCA attention, manifold-constrained hyper-connections, and three-tier reasoning (Non-think / Think High / Think Max). - [deepseek-ai/DeepSeek-V3.1](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V3.1): DeepSeek-V3.1 is a hybrid MoE model that supports dynamic switching between thinking and non-thinking modes, with tool calling and function execution. - [deepseek-ai/DeepSeek-V3.2-Exp](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V3.2-Exp): Experimental DeepSeek-V3.2 preview with sparse attention (MQA-like logits) and FP8 KV cache; architecture matches DeepSeek-V3.1 except for the sparse attention mechanism. - [deepseek-ai/DeepSeek-V3.2](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V3.2): DeepSeek V3.2 MoE model with MLA attention, sparse attention, and scalable RL for strong reasoning and agent capabilities. - [deepseek-ai/DeepSeek-V4-Flash-Vision-Exp](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp): DeepSeek's first experimental multimodal V4 model — the V4-Flash MoE backbone plus a 32-layer vision tower, 1M context, and a fused DSpark draft module. - [deepseek-ai/DeepSeek-V3](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V3): DeepSeek-V3 is a 671B-parameter Mixture-of-Experts model with native FP8 weights and strong reasoning, coding, and math capabilities. - [deepseek-ai/DeepSeek-OCR-2](https://recipes.vllm.ai/deepseek-ai/DeepSeek-OCR-2): Next-generation DeepSeek OCR model with improved document-to-markdown grounding and optical context compression. - [deepseek-ai/DeepSeek-OCR](https://recipes.vllm.ai/deepseek-ai/DeepSeek-OCR): Frontier OCR model exploring optical context compression for LLMs, optimized for document parsing and markdown generation. ### Qwen - [Qwen/Qwen3.8-27B](https://recipes.vllm.ai/Qwen/Qwen3.8-27B): 27B-parameter dense hybrid-attention model with linear attention on 48 of 64 layers, a vision tower, a built-in MTP draft head, 262K native context window and extensible to 1M context - [Qwen/Qwen3.8-Flash-Next](https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next): Qwen4 architecture preview with a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. - [Qwen/Qwen3-30B-A3B](https://recipes.vllm.ai/Qwen/Qwen3-30B-A3B): Qwen3 MoE model with 30.5B total and 3.3B active parameters, validated for BF16 serving on Intel Xeon 6 and Intel Arc Pro B70 and W8A8 serving on Ascend A3. - [Qwen/Qwen-Image-2.1](https://recipes.vllm.ai/Qwen/Qwen-Image-2.1): Unified text-to-image and image-conditioned generation. A 7.1B single-stream DiT with block-causal attention and an exact cross-step prefix KV cache, paired with a Qwen3-VL-8B text encoder and a 16x RGBA autoencoder. - [Qwen/Qwen3.5-122B-A10B](https://recipes.vllm.ai/Qwen/Qwen3.5-122B-A10B): Mid-size Qwen3.5 multimodal MoE (122B total / 10B active) with gated delta networks, 256 experts, and 262K context - [Qwen/Qwen3.6-27B](https://recipes.vllm.ai/Qwen/Qwen3.6-27B): Qwen3.6 dense multimodal model (27B) with gated delta networks hybrid attention, MTP, and 262K context - [Qwen/Qwen3.5-0.8B](https://recipes.vllm.ai/Qwen/Qwen3.5-0.8B): Qwen3.5 tiny dense multimodal model (0.8B) — ultra-low-VRAM / edge serving with 262K context - [Qwen/Qwen3.5-27B](https://recipes.vllm.ai/Qwen/Qwen3.5-27B): Qwen3.5 dense multimodal model (27B) with gated delta networks hybrid attention, MTP, and 262K context - [Qwen/Qwen3.5-2B](https://recipes.vllm.ai/Qwen/Qwen3.5-2B): Qwen3.5 mini dense multimodal model (2B) — edge / low-VRAM serving with 262K context - [Qwen/Qwen3.5-35B-A3B](https://recipes.vllm.ai/Qwen/Qwen3.5-35B-A3B): Compact Qwen3.5 multimodal MoE (35B total / 3B active) with gated delta networks, 256 experts, and 262K context - [Qwen/Qwen3.5-397B-A17B](https://recipes.vllm.ai/Qwen/Qwen3.5-397B-A17B): Multimodal MoE model with gated delta networks architecture, 397B total / 17B active parameters, up to 262K context - [Qwen/Qwen3.5-4B](https://recipes.vllm.ai/Qwen/Qwen3.5-4B): Qwen3.5 compact dense multimodal model (4B) — fits on 16 GB consumer GPUs with full 262K context or one Xeon 6 NUMA node or Intel Arc Pro B60/B70 - [Qwen/Qwen3.5-9B](https://recipes.vllm.ai/Qwen/Qwen3.5-9B): Qwen3.5 dense multimodal model (9B) with gated delta networks hybrid attention, MTP, and 262K context - [Qwen/Qwen3.6-35B-A3B](https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B): Smaller Qwen3.6 multimodal MoE model (35B total / 3B active) with BF16, FP8, and NVIDIA NVFP4 variants - [Qwen/Qwen3.8-2.4T-A95B](https://recipes.vllm.ai/Qwen/Qwen3.8-2.4T-A95B): 2.4T-parameter hybrid-attention MoE (~95B active) with linear attention on 69 of 92 layers, 512 routed experts, a built-in MTP draft head, 262K native context window and extensible to 1M context - [Qwen/QwQ-32B](https://recipes.vllm.ai/Qwen/QwQ-32B): Qwen's 32B dense reasoning model, validated for Intel Xeon 6 CPU execution. - [Qwen/Qwen2.5-VL-7B-Instruct](https://recipes.vllm.ai/Qwen/Qwen2.5-VL-7B-Instruct): Qwen2.5-VL dense vision-language model (7B) for image and video understanding — fits on a single TPU v6e chip, one GPU, or Intel Xeon 6 CPU. - [Qwen/Qwen3-1.7B](https://recipes.vllm.ai/Qwen/Qwen3-1.7B): Qwen3 1.7B dense model with hybrid thinking/non-thinking modes, validated for BF16 serving on Intel Xeon 6. - [Qwen/Qwen3-14B](https://recipes.vllm.ai/Qwen/Qwen3-14B): Qwen3 14B dense model with hybrid thinking/non-thinking modes, validated for Intel Xeon 6 CPU serving. - [Qwen/Qwen3-4B](https://recipes.vllm.ai/Qwen/Qwen3-4B): Qwen3 4B dense model with hybrid thinking/non-thinking modes — fits on a single TPU v6e chip, one GPU or one Xeon 6 NUMA node. - [Qwen/Qwen3-8B](https://recipes.vllm.ai/Qwen/Qwen3-8B): Qwen3 8.2B dense model with hybrid thinking/non-thinking modes, validated for BF16 serving on Intel Xeon 6. - [Qwen/Qwen3-VL-30B-A3B-Instruct](https://recipes.vllm.ai/Qwen/Qwen3-VL-30B-A3B-Instruct): Qwen3-VL MoE vision-language model with 30B total / 3B active parameters, supporting image, video, and text workloads. - [Qwen/Qwen3-32B](https://recipes.vllm.ai/Qwen/Qwen3-32B): Qwen3 32B dense model with hybrid thinking/non-thinking modes — verified on TPU v6e (Trillium). - [Qwen/Qwen3-Next-80B-A3B-Instruct](https://recipes.vllm.ai/Qwen/Qwen3-Next-80B-A3B-Instruct): Advanced Qwen3-Next MoE model (80B total / 3B active) with hybrid attention, highly sparse experts, and multi-token prediction. - [Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice](https://recipes.vllm.ai/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice): Text-to-speech served via vLLM-Omni with predefined speaker voices and optional style/emotion control, exposed through the OpenAI /v1/audio/speech API. Sibling VoiceDesign and Base (voice-clone) checkpoints share the same serving path. - [Qwen/Qwen2.5-32B](https://recipes.vllm.ai/Qwen/Qwen2.5-32B): Qwen2.5 32B dense base (pretrained) language model for text completion — verified on TPU v6e (Trillium). - [Qwen/Qwen3Guard-Gen-8B](https://recipes.vllm.ai/Qwen/Qwen3Guard-Gen-8B): Lightweight text-only guardrail/safety classifier model in the Qwen3Guard family. - [Qwen/Qwen-Image](https://recipes.vllm.ai/Qwen/Qwen-Image): Text-to-image diffusion model (20B parameters) from the Qwen-Image family, served via vLLM-Omni. - [Qwen/Qwen2.5-VL-72B-Instruct](https://recipes.vllm.ai/Qwen/Qwen2.5-VL-72B-Instruct): Qwen2.5-VL dense vision-language model (72B) for high-quality image and video understanding. - [Qwen/Qwen3-235B-A22B-Instruct-2507](https://recipes.vllm.ai/Qwen/Qwen3-235B-A22B-Instruct-2507): Flagship Qwen3 MoE instruct model with 235B total and 22B active parameters, tuned for high-quality text generation. - [Qwen/Qwen3-ASR-1.7B](https://recipes.vllm.ai/Qwen/Qwen3-ASR-1.7B): Speech-to-text model supporting 11 languages, multiple accents, and singing voice with customizable text-context prompting. - [Qwen/Qwen3-Coder-480B-A35B-Instruct](https://recipes.vllm.ai/Qwen/Qwen3-Coder-480B-A35B-Instruct): Large coder MoE with 480B total / 35B active parameters, strong tool-use and code generation capabilities. - [Qwen/Qwen3-VL-235B-A22B-Instruct](https://recipes.vllm.ai/Qwen/Qwen3-VL-235B-A22B-Instruct): Qwen3-VL flagship MoE vision-language model with 235B total / 22B active parameters, supporting images, video, and long context. ### Thinking Machines Lab - [thinkingmachines/Inkling-Small](https://recipes.vllm.ai/thinkingmachines/Inkling-Small): Natively multimodal 276B-parameter MoE from Thinking Machines Lab — 12B active parameters, text/image/audio in, text out, and up to 1M context. - [thinkingmachines/Inkling](https://recipes.vllm.ai/thinkingmachines/Inkling): Natively multimodal 1T-parameter MoE from Thinking Machines Lab — text/image/audio in, text out, up to 1M context — with relative attention, short convolution, and shared expert sinks. ### Hunyuan (Tencent) - [tencent/Hy4-preview](https://recipes.vllm.ai/tencent/Hy4-preview): Tencent Hunyuan Hy4-preview — scaled-up MoE language model (770B total / 49B active) with a 10B MTP layer for speculative decoding, 1M context, and hy_v4 tool/reasoning parsers - [tencent/Hy3-preview](https://recipes.vllm.ai/tencent/Hy3-preview): Tencent Hunyuan Hy3-preview — scaled-up MoE language model (295B total / 21B active) with a 3.8B MTP layer for speculative decoding, 256K context, and hy_v3 tool/reasoning parsers - [tencent/Hy3](https://recipes.vllm.ai/tencent/Hy3): Tencent Hy3 — scaled-up MoE language model (295B total / 21B active) with a 3.8B MTP layer for speculative decoding, 256K context, and hy_v3 tool/reasoning parsers - [tencent/Hunyuan-A13B-Instruct](https://recipes.vllm.ai/tencent/Hunyuan-A13B-Instruct): Tencent Hunyuan A13B instruct-tuned MoE language model with AITER-accelerated AMD ROCm deployment - [tencent/HunyuanOCR](https://recipes.vllm.ai/tencent/HunyuanOCR): Tencent Hunyuan end-to-end OCR expert VLM (~1B) for online OCR serving with an OpenAI-compatible API ### Moonshot AI - [moonshotai/Kimi-K3](https://recipes.vllm.ai/moonshotai/Kimi-K3): Pre-release 2.8T-parameter native multimodal MoE with Kimi Delta Attention, Gated MLA, Attention Residuals, and a 1M-token context window - [moonshotai/Kimi-K2.5](https://recipes.vllm.ai/moonshotai/Kimi-K2.5): Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes - [moonshotai/Kimi-K2.6](https://recipes.vllm.ai/moonshotai/Kimi-K2.6): Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes - [moonshotai/Kimi-K2.7-Code](https://recipes.vllm.ai/moonshotai/Kimi-K2.7-Code): Coding-focused agentic MoE built on Kimi-K2.6, tuned for long-horizon software-engineering tasks with thinking-only reasoning, tool calling, and vision-language input - [moonshotai/Kimi-K2-Instruct](https://recipes.vllm.ai/moonshotai/Kimi-K2-Instruct): Moonshot AI's Kimi-K2 is a trillion-parameter MoE instruction model (~32B active) with native FP8 weights and strong tool-calling capabilities. - [moonshotai/Kimi-K2-Thinking](https://recipes.vllm.ai/moonshotai/Kimi-K2-Thinking): Kimi-K2-Thinking is an advanced reasoning MoE model with native INT4 QAT weights, designed for long-horizon agent workflows interleaving chain-of-thought reasoning with tool calls. - [moonshotai/Kimi-Linear-48B-A3B-Instruct](https://recipes.vllm.ai/moonshotai/Kimi-Linear-48B-A3B-Instruct): Kimi-Linear is a 48B-parameter instruction-tuned MoE model (~3B activated) with a linear-attention variant supporting very long context (1M tokens). ### Google - [Google/diffusiongemma-26B-A4B-it](https://recipes.vllm.ai/Google/diffusiongemma-26B-A4B-it): Google's DiffusionGemma — a block-diffusion language model built on Gemma 4's MoE backbone (26B total / 4B active). Generates tokens via iterative denoising over a fixed-length canvas rather than left-to-right autoregressive decoding, enabling higher throughput with parallel block generation. - [Google/gemma-4-26B-A4B-it](https://recipes.vllm.ai/Google/gemma-4-26B-A4B-it): Google's Gemma 4 MoE multimodal model (26B total / 4B active) with 128 fine-grained experts, top-8 routing, thinking mode, and tool-use protocol. - [Google/gemma-4-E2B-it](https://recipes.vllm.ai/Google/gemma-4-E2B-it): Google's compact Gemma 4 multimodal model (effective 2B) with native text, image, and audio, plus thinking mode and tool-use protocol. - [Google/gemma-4-E4B-it](https://recipes.vllm.ai/Google/gemma-4-E4B-it): Google's compact Gemma 4 multimodal model (effective 4B) with native text, image, and audio, plus thinking mode and tool-use protocol. - [Google/gemma-4-12B-it](https://recipes.vllm.ai/Google/gemma-4-12B-it): Google's encoder-free unified Gemma 4 dense model (12B) with native text, image, and audio, plus thinking mode and tool-use protocol. - [Google/gemma-4-31B-it](https://recipes.vllm.ai/Google/gemma-4-31B-it): Google's unified multimodal Gemma 4 dense model (31B) with native text, image, and audio, plus thinking mode and tool-use protocol. - [Google/translategemma-27b-it](https://recipes.vllm.ai/Google/translategemma-27b-it): Lightweight open translation model from Google (based on Gemma 3) supporting 55 languages. Served via the vLLM-optimized Infomaniak-AI checkpoint. ### Meta - [meta-llama/Llama-3.3-70B-Instruct](https://recipes.vllm.ai/meta-llama/Llama-3.3-70B-Instruct): Llama 3.3 70B dense model with NVIDIA FP8/FP4 quantized variants for Hopper and Blackwell GPUs - [meta-llama/Llama-3.1-8B-Instruct](https://recipes.vllm.ai/meta-llama/Llama-3.1-8B-Instruct): Meta's Llama 3.1 8B dense instruction-tuned language model with 128K context - [meta-llama/Llama-3.1-8B](https://recipes.vllm.ai/meta-llama/Llama-3.1-8B): Meta Llama 3.1 8B dense base language model with a CPU-validated Red Hat AI W8A8 variant. - [meta-llama/Llama-3.2-1B-Instruct](https://recipes.vllm.ai/meta-llama/Llama-3.2-1B-Instruct): Meta Llama 3.2 1B Instruct with CPU-validated Red Hat AI W8A8 and AMead10 AWQ variants. - [meta-llama/Llama-3.2-1B](https://recipes.vllm.ai/meta-llama/Llama-3.2-1B): Meta Llama 3.2 1.23B dense pretrained language model with 128K context, validated for BF16 serving on Intel Xeon 6. - [meta-llama/Llama-3.2-3B-Instruct](https://recipes.vllm.ai/meta-llama/Llama-3.2-3B-Instruct): Meta Llama 3.2 3.21B dense instruction-tuned language model with 128K context, validated for BF16 serving on Intel Xeon 6. - [meta-llama/Llama-4-Scout-17B-16E-Instruct](https://recipes.vllm.ai/meta-llama/Llama-4-Scout-17B-16E-Instruct): Llama 4 Scout 17B-16E MoE model with NVIDIA FP8/FP4 variants, fits on a single GPU with quantization ### NVIDIA - [nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16): NVIDIA Nemotron 3.5 Lightning hybrid Mamba-MoE (30B total / 3B active) with NVFP4 and BF16 checkpoints, 1M context, and MTP / DSpark / DFlash speculative decoding - [nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16): NVIDIA Nemotron-3-Nano Mamba-hybrid MoE (30B total / ~3B active) with BF16 and FP8 variants - [nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16): NVIDIA Nemotron-3-Super Mamba-hybrid latent-MoE (~120B total / ~12B active) with BF16, FP8, and NVFP4 variants - [nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16): NVIDIA Nemotron 3 Ultra hybrid Transformer-Mamba MoE model for long-context agentic reasoning, coding, and tool use. - [nvidia/Cosmos3-Nano](https://recipes.vllm.ai/nvidia/Cosmos3-Nano): Compact 16B omnimodal world model (Mixture-of-Transformers) for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI - [nvidia/Cosmos3-Super-Image2Video](https://recipes.vllm.ai/nvidia/Cosmos3-Super-Image2Video): 64B Cosmos3-Super specialization for temporally coherent image-to-video generation - [nvidia/Cosmos3-Super-Text2Image](https://recipes.vllm.ai/nvidia/Cosmos3-Super-Text2Image): 64B Cosmos3-Super specialization for high-fidelity text-to-image generation - [nvidia/Cosmos3-Super](https://recipes.vllm.ai/nvidia/Cosmos3-Super): Frontier-scale 64B omnimodal world model (Mixture-of-Transformers) for advanced multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI - [nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16](https://recipes.vllm.ai/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16): Mamba2-Transformer hybrid MoE omnimodal model (31B total / 3B active) with unified video, audio, image, and text understanding; reasoning + tool calling; BF16, FP8, and NVFP4 variants - [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16): NVIDIA Nemotron-3-Nano 4B (Mamba-hybrid dense) — compact reasoning + tool-use model with BF16 and FP8 variants - [nvidia/NVIDIA-Nemotron-Nano-9B-v2](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-Nano-9B-v2): NVIDIA Nemotron-Nano 9B (Mamba-hybrid dense) reasoning + tool-use model with FP8 / NVFP4 / Japanese variants - [nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16](https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16): NVIDIA Nemotron-Nano 12B vision-language model with video support and Efficient Video Sampling (EVS) ### OpenAI - [openai/gpt-oss-120b](https://recipes.vllm.ai/openai/gpt-oss-120b): OpenAI's gpt-oss family (20B / 120B) with MXFP4 MoE, attention-sinks, built-in tools via Responses API - [openai/gpt-oss-20b](https://recipes.vllm.ai/openai/gpt-oss-20b): OpenAI's gpt-oss-20b — 21B-total / 3.6B-active MoE reasoning model with native MXFP4 quant; supports GPU, XPU, and Xeon 6 CPU serving - [openai/whisper-large-v3](https://recipes.vllm.ai/openai/whisper-large-v3): Whisper Large v3 speech recognition model, validated for Intel Xeon 6 CPU serving. ### BharatGen - [bharatgenai/Param2-17B-A2.4B-Thinking](https://recipes.vllm.ai/bharatgenai/Param2-17B-A2.4B-Thinking): BharatGen's hybrid-MoE reasoning model for English, Hindi and 21 Indian languages, with thinking traces and Hermes-style tool calling ### InternLM - [internlm/Intern-S2-397B](https://recipes.vllm.ai/internlm/Intern-S2-397B): Flagship scientific multimodal MoE (397B total / 17B active) on the Qwen3.5 hybrid linear/full-attention architecture — 262K context, MTP-accelerated reasoning. BF16 and FP8 checkpoints. - [internlm/Intern-S2-Preview](https://recipes.vllm.ai/internlm/Intern-S2-Preview): Scientific multimodal MoE (36B total / 3B active) continued pre-trained from Qwen3.5 — hybrid linear/full attention, 262K context, MTP-accelerated reasoning. BF16 and FP8 checkpoints. - [internlm/Intern-S1](https://recipes.vllm.ai/internlm/Intern-S1): Intern-S1 vision-language model from Shanghai AI Lab with BF16/FP8 variants and thinking/non-thinking modes ### Preferred Networks - [pfnet/plamo-3-nict-31b-base](https://recipes.vllm.ai/pfnet/plamo-3-nict-31b-base): Largest PLaMo 3 NICT Japanese/English base model with interleaved sliding-window and full attention. ### Microsoft - [microsoft/Phi-4-multimodal-instruct](https://recipes.vllm.ai/microsoft/Phi-4-multimodal-instruct): Microsoft Phi-4 multimodal model supporting text, image, and audio inputs with text outputs. - [microsoft/Phi-4-reasoning](https://recipes.vllm.ai/microsoft/Phi-4-reasoning): Microsoft Phi-4 14B dense reasoning model with 32K context. - [microsoft/Phi-4-mini-instruct](https://recipes.vllm.ai/microsoft/Phi-4-mini-instruct): Microsoft's Phi-4 family of lightweight dense models (mini-instruct, reasoning, multimodal) with 128K context ### Dots - [dots-studio/dots3-note-prev](https://recipes.vllm.ai/dots-studio/dots3-note-prev): Multimodal MoE model available in BF16 and native FP8, with hybrid DSA and SWA, 512K context, and MTP speculative decoding. ### inclusionAI - [inclusionAI/Ling-3.0-flash](https://recipes.vllm.ai/inclusionAI/Ling-3.0-flash): Ling-3.0-flash MoE model with BF16, FP8, FP4, and INT4 checkpoints, native MTP, and an external DSpark draft model - [inclusionAI/Ling-3.0-tiny](https://recipes.vllm.ai/inclusionAI/Ling-3.0-tiny): Ling-3.0-tiny lightweight hybrid-reasoning MoE with BF16, block-FP8, and compressed-tensors INT4 checkpoints - [inclusionAI/Ling-3.0-flash-VL](https://recipes.vllm.ai/inclusionAI/Ling-3.0-flash-VL): Ling-3.0-flash vision-language MoE model with BF16, FP8, FP4, and INT4 checkpoints - [inclusionAI/Ming-omni-tts-0.5B](https://recipes.vllm.ai/inclusionAI/Ming-omni-tts-0.5B): inclusionAI's dense 0.5B-LM text-to-speech model served via vLLM-Omni with style, dialect, voice-cloning and multi-speaker controls, through the OpenAI /v1/audio/speech API (44.1 kHz mono). - [inclusionAI/Ring-2.6-1T](https://recipes.vllm.ai/inclusionAI/Ring-2.6-1T): Ring-2.6-1T (BailingMoeV2_5) FP8 thinking model with 1T total / 50B active params, hybrid linear + MLA attention, 128K context - [inclusionAI/Ling-2.6-1T](https://recipes.vllm.ai/inclusionAI/Ling-2.6-1T): Ling-2.6-1T (BailingMoeV2_5) FP8 instruct model with 1T total / 50B active params, hybrid linear + MLA attention, 262K context - [inclusionAI/Ling-2.6-flash](https://recipes.vllm.ai/inclusionAI/Ling-2.6-flash): Ling-2.6-flash (BailingMoeV2_5) instruct model with 104B total / 7.4B active params, hybrid linear + MLA attention, 128K context, optimized for agent workloads - [inclusionAI/Ring-1T-FP8](https://recipes.vllm.ai/inclusionAI/Ring-1T-FP8): Ring-1T (BailingMoeV2) FP8 model (~1T total params) for 8xH200 or 8xMI300X deployment ### Muse (Meta) - [meta-models/Muse-Glimmer-30B](https://recipes.vllm.ai/meta-models/Muse-Glimmer-30B): Dense 29.6B vision-language model with a ViT-G/14 perception encoder and 128K context, distilled from Muse Spark for local agentic use. Emits channel-scoped reasoning and XML-style ATEM tool calls rather than JSON, so it needs the dedicated `muse_glimmer` tool-call and reasoning parsers. ### IFM - [IFM/K2-Horizon-375B-A23B](https://recipes.vllm.ai/IFM/K2-Horizon-375B-A23B): Frontier-scale sparse MoE model for long-context research and high-capacity serving - [IFM/K2-Horizon-MoVA-36B-A4B](https://recipes.vllm.ai/IFM/K2-Horizon-MoVA-36B-A4B): Sparse MoVA and MoE long-context model for research and production-style serving experiments - [IFM/K2-Horizon-3.7B](https://recipes.vllm.ai/IFM/K2-Horizon-3.7B): Dense model for long-context research, serving, and downstream adaptation - [IFM/K2-Horizon-32B](https://recipes.vllm.ai/IFM/K2-Horizon-32B): Dense long-context model for reasoning research and high-capacity serving - [IFM/K2-Horizon-7B](https://recipes.vllm.ai/IFM/K2-Horizon-7B): Dense model for long-context research, fine-tuning, and cost-conscious deployment - [IFM/K2-Horizon-0.9B](https://recipes.vllm.ai/IFM/K2-Horizon-0.9B): Compact dense model for mathematics, code, instruction following, and STEM tasks ### MiniCPM (OpenBMB) - [openbmb/MiniCPM5-2B](https://recipes.vllm.ai/openbmb/MiniCPM5-2B): MiniCPM5-2B — dense 2B LLM with hybrid Think/No-Think reasoning, native 128K context, and native tool calling support, built on the standard Llama architecture - [openbmb/VoxCPM2](https://recipes.vllm.ai/openbmb/VoxCPM2): OpenBMB's 2B native-AR text-to-speech model served via vLLM-Omni — 48 kHz mono, 30+ languages, zero-shot synthesis and reference-audio voice cloning — through the OpenAI /v1/audio/speech API. - [openbmb/MiniCPM-V-4.6](https://recipes.vllm.ai/openbmb/MiniCPM-V-4.6): MiniCPM-V 4.6 (1.3B) — pocket-sized multimodal LLM for ultra-efficient single-image, multi-image, and video understanding, built on SigLIP2-400M + a Qwen3.5-0.8B hybrid-attention backbone - [openbmb/MiniCPM5-1B](https://recipes.vllm.ai/openbmb/MiniCPM5-1B): MiniCPM5-1B — dense 1B on-device LLM with hybrid Think/No-Think reasoning, native 128K context, and strong agentic tool use, built on the standard Llama architecture ### ibm-granite - [ibm-granite/granite-3.2-2b-instruct](https://recipes.vllm.ai/ibm-granite/granite-3.2-2b-instruct): IBM Granite 3.2 2B instruction-tuned dense language model with long-context support. ### Lightricks - [Lightricks/LTX-2.5-Diffusers](https://recipes.vllm.ai/Lightricks/LTX-2.5-Diffusers): 19B diffusion transformer for joint video and synchronized audio generation, served via vLLM-Omni ### Stability AI - [stabilityai/stable-audio-open-1.0](https://recipes.vllm.ai/stabilityai/stable-audio-open-1.0): Text-to-audio generation model (1.2B params) producing up to ~47 s stereo audio at 44.1 kHz, served via vLLM-Omni - [stabilityai/stable-diffusion-3.5-medium](https://recipes.vllm.ai/stabilityai/stable-diffusion-3.5-medium): Stability AI's Stable Diffusion 3.5 text-to-image family (medium 2.5B, large 8.1B, large-turbo) via vLLM-Omni with Cache-DiT acceleration ### IndexTeam - [IndexTeam/IndexTTS-2.5](https://recipes.vllm.ai/IndexTeam/IndexTTS-2.5): Multilingual zero-shot voice-cloning TTS with native speed, emotion, and text-normalization controls, served through vLLM-Omni's OpenAI-compatible speech API. ### OpenMOSS - [OpenMOSS-Team/MOSS-Transcribe-Diarize](https://recipes.vllm.ai/OpenMOSS-Team/MOSS-Transcribe-Diarize): OpenMOSS's 0.9B end-to-end multi-speaker long-audio transcription model with timestamps and speaker labels, served through vLLM's OpenAI-compatible /v1/audio/transcriptions API. - [OpenMOSS-Team/MOSS-SoundEffect](https://recipes.vllm.ai/OpenMOSS-Team/MOSS-SoundEffect): OpenMOSS's 8B sound-effect generation model — environmental, urban, biological, human-action and musical sounds with controllable duration, no reference audio — served via vLLM-Omni through the OpenAI /v1/audio/speech API (24 kHz mono). - [OpenMOSS-Team/MOSS-TTS-Realtime](https://recipes.vllm.ai/OpenMOSS-Team/MOSS-TTS-Realtime): OpenMOSS's 1.7B real-time streaming TTS for low-latency voice agents (TTFB ~180 ms) — multi-turn context-aware incremental synthesis — served via vLLM-Omni through the OpenAI /v1/audio/speech API (24 kHz mono). - [OpenMOSS-Team/MOSS-TTS](https://recipes.vllm.ai/OpenMOSS-Team/MOSS-TTS): Flagship 8B model of the OpenMOSS MOSS-TTS Family — high-fidelity zero-shot voice cloning, ultra-long stable speech, token-level duration and phoneme control, 20-language code-switched synthesis — served via vLLM-Omni through the OpenAI /v1/audio/speech API (24 kHz mono). - [OpenMOSS-Team/MOSS-TTSD-v1.0](https://recipes.vllm.ai/OpenMOSS-Team/MOSS-TTSD-v1.0): OpenMOSS's 8B spoken-dialogue generation model for expressive, multi-speaker, ultra-long conversations — served via vLLM-Omni through the OpenAI /v1/audio/speech API (24 kHz mono). - [OpenMOSS-Team/MOSS-VoiceGenerator](https://recipes.vllm.ai/OpenMOSS-Team/MOSS-VoiceGenerator): OpenMOSS's 1.7B zero-shot voice-design model — generate diverse voices and styles directly from a text prompt with no reference speech — served via vLLM-Omni through the OpenAI /v1/audio/speech API (24 kHz mono). ### MindLab Research - [mindlab-research/Macaron-V1-Coding-Venti](https://recipes.vllm.ai/mindlab-research/Macaron-V1-Coding-Venti): Macaron-V1-Coding-Venti — MindLab's coding-specialist checkpoint: the Macaron-V1-Venti L2 Coding LoRA merged into the GLM-5.2 BF16 base. Same MoE architecture and launch as GLM-5.2 (~743B total, 39B active), no runtime adapter. ### Poolside - [poolside/Laguna-S-2.1](https://recipes.vllm.ai/poolside/Laguna-S-2.1): Poolside's 118B total / 8B activated MoE coding model with mixed sliding-window + global attention, native interleaved reasoning, and 256K context — the larger sibling of Laguna XS-2.1, tuned for agentic coding. - [poolside/Laguna-XS-2.1](https://recipes.vllm.ai/poolside/Laguna-XS-2.1): Poolside's 33B total / 3B activated MoE coding model with mixed sliding-window + global attention, native interleaved reasoning, and 256K context — designed for agentic coding. - [poolside/Laguna-M.1](https://recipes.vllm.ai/poolside/Laguna-M.1): Poolside's 225B total / 23B activated MoE coding model with global attention, native interleaved reasoning, and 256K context — designed for agentic coding and long-horizon work. - [poolside/Laguna-XS.2](https://recipes.vllm.ai/poolside/Laguna-XS.2): Poolside's 33B total / 3B activated MoE coding model with mixed sliding-window + global attention, native interleaved reasoning, and 256K context — designed for agentic coding. ### PaddlePaddle - [PaddlePaddle/PaddleOCR-VL-1.6](https://recipes.vllm.ai/PaddlePaddle/PaddleOCR-VL-1.6): PaddleOCR-VL-1.6 (0.9B) — region-aware data optimization + progressive post-training; new SOTA 96.33% on OmniDocBench v1.6, drop-in replacement for 1.5 - [PaddlePaddle/PaddleOCR-VL-1.5](https://recipes.vllm.ai/PaddlePaddle/PaddleOCR-VL-1.5): PaddleOCR-VL-1.5 (0.9B) — next-gen compact VLM for document parsing; adds text spotting, seal recognition, and Tibetan/Bengali - [PaddlePaddle/PaddleOCR-VL](https://recipes.vllm.ai/PaddlePaddle/PaddleOCR-VL): PaddleOCR-VL (0.9B) — compact vision-language model for document parsing, OCR, tables, formulas, charts ### Ernie (Baidu) - [baidu/Unlimited-OCR](https://recipes.vllm.ai/baidu/Unlimited-OCR): Baidu's state-of-the-art document-parsing model with Reference Sliding Window Attention (R-SWA), optimized for full-page OCR and markdown generation. - [baidu/ERNIE-4.5-21B-A3B-PT](https://recipes.vllm.ai/baidu/ERNIE-4.5-21B-A3B-PT): Baidu ERNIE 4.5 MoE text models (21B-A3B, 300B-A47B) with BF16 and FP8 support plus ERNIE-MTP speculative decoding - [baidu/ERNIE-4.5-VL-28B-A3B-PT](https://recipes.vllm.ai/baidu/ERNIE-4.5-VL-28B-A3B-PT): Baidu ERNIE 4.5 VL MoE vision-language models (28B-A3B, 424B-A47B) with heterogeneous text/vision experts ### Liquid AI - [LiquidAI/LFM2.5-230M](https://recipes.vllm.ai/LiquidAI/LFM2.5-230M): Liquid AI's most compact LFM2.5 chat model (230M) on the LFM2 hybrid conv+attention backbone — distilled from LFM2.5-350M, tool calling and 32K context, built for edge / on-device serving. - [LiquidAI/LFM2.5-1.2B-Base](https://recipes.vllm.ai/LiquidAI/LFM2.5-1.2B-Base): Liquid AI's 1.2B pretrained base model on the LFM2 hybrid conv+attention backbone — a text-completion and fine-tuning foundation (no chat template). - [LiquidAI/LFM2.5-1.2B-Instruct](https://recipes.vllm.ai/LiquidAI/LFM2.5-1.2B-Instruct): Liquid AI's 1.2B instruction-tuned model on the LFM2 hybrid conv+attention backbone, with tool calling and a 32K context window on a single small GPU. - [LiquidAI/LFM2.5-1.2B-JP-202606](https://recipes.vllm.ai/LiquidAI/LFM2.5-1.2B-JP-202606): Liquid AI's updated (2026-06) 1.2B Japanese chat model on the LFM2 hybrid conv+attention backbone — adds tool calling, with a 32K context window. - [LiquidAI/LFM2.5-1.2B-JP](https://recipes.vllm.ai/LiquidAI/LFM2.5-1.2B-JP): Liquid AI's 1.2B Japanese-specialized chat model on the LFM2 hybrid conv+attention backbone, with a 32K context window on a single small GPU. - [LiquidAI/LFM2.5-1.2B-Thinking](https://recipes.vllm.ai/LiquidAI/LFM2.5-1.2B-Thinking): Liquid AI's 1.2B reasoning model on the LFM2 hybrid conv+attention backbone — chain-of-thought, tool calling, and 32K context on a single small GPU. - [LiquidAI/LFM2.5-350M](https://recipes.vllm.ai/LiquidAI/LFM2.5-350M): Liquid AI's smallest LFM2.5 chat model (350M) on the LFM2 hybrid conv+attention backbone — tool calling and 32K context, light enough for edge GPUs. - [LiquidAI/LFM2.5-8B-A1B](https://recipes.vllm.ai/LiquidAI/LFM2.5-8B-A1B): Liquid AI's 8B mixture-of-experts model (~1B active) on the LFM2 hybrid conv+attention backbone, with reasoning and tool calling at ~1B decode cost. - [LiquidAI/LFM2.5-VL-1.6B](https://recipes.vllm.ai/LiquidAI/LFM2.5-VL-1.6B): Liquid AI's 1.6B vision-language model — LFM2 hybrid LM backbone plus a SigLIP2 vision tower for image+text chat on a single small GPU. - [LiquidAI/LFM2.5-VL-450M](https://recipes.vllm.ai/LiquidAI/LFM2.5-VL-450M): Liquid AI's smallest vision-language model (450M) — LFM2 hybrid LM backbone plus a SigLIP2 vision tower for image+text chat, light enough for edge GPUs. ### Boson AI - [bosonai/higgs-audio-v3-tts-4b](https://recipes.vllm.ai/bosonai/higgs-audio-v3-tts-4b): Boson AI's ~4B Qwen3-backbone text-to-speech model served via vLLM-Omni — 24 kHz speech, 100+ languages, zero-shot voice cloning with inline emotion/style/prosody control tokens — through the OpenAI /v1/audio/speech API. ### Fish Audio - [fishaudio/s2-pro](https://recipes.vllm.ai/fishaudio/s2-pro): Fish Audio's dual-AR text-to-speech model served via vLLM-Omni, producing 44.1 kHz mono audio with optional voice cloning, through the OpenAI /v1/audio/speech API. ### Mistral AI - [mistralai/Voxtral-4B-TTS-2603](https://recipes.vllm.ai/mistralai/Voxtral-4B-TTS-2603): Mistral's 4B text-to-speech model served via vLLM-Omni with built-in voice presets, exposed through the OpenAI /v1/audio/speech API (24 kHz mono). - [mistralai/Mistral-Medium-3.5-128B](https://recipes.vllm.ai/mistralai/Mistral-Medium-3.5-128B): Mistral Medium 3.5 (128B) dense vision-language model with native FP8 weights and 256K context - [mistralai/Ministral-3-14B-Instruct-2512](https://recipes.vllm.ai/mistralai/Ministral-3-14B-Instruct-2512): Ministral 3 Instruct family (3B/8B/14B) with FP8 weights, vision support, and 256K context - [mistralai/Mistral-Small-4-119B-2603](https://recipes.vllm.ai/mistralai/Mistral-Small-4-119B-2603): Mistral Small 4 (119B MoE, 6.5B active) — multimodal hybrid instruct + reasoning model with native FP8 weights and 256K context - [mistralai/Voxtral-Mini-4B-Realtime-2602](https://recipes.vllm.ai/mistralai/Voxtral-Mini-4B-Realtime-2602): Multilingual realtime speech transcription (13 languages) with a natively streaming causal audio encoder; configurable 80ms–2.4s transcription delay served via vLLM's Realtime API - [mistralai/Ministral-3-8B-Reasoning-2512](https://recipes.vllm.ai/mistralai/Ministral-3-8B-Reasoning-2512): Ministral 3 Reasoning family (3B/8B/14B) with BF16 weights, vision support, and 256K context - [mistralai/Mistral-Large-3-675B-Instruct-2512](https://recipes.vllm.ai/mistralai/Mistral-Large-3-675B-Instruct-2512): Mistral Large 3 (675B) with FP8 and NVFP4 weights for 8xH200 / 4xB200 deployments ### JetBrains - [JetBrains/Mellum2-12B-A2.5B-Instruct](https://recipes.vllm.ai/JetBrains/Mellum2-12B-A2.5B-Instruct): JetBrains' instruction-tuned code MoE (12B total / 2.5B active) that answers directly without an externalized chain of thought — low-latency coding and tool use - [JetBrains/Mellum2-12B-A2.5B-Thinking](https://recipes.vllm.ai/JetBrains/Mellum2-12B-A2.5B-Thinking): JetBrains' reasoning-augmented code MoE (12B total / 2.5B active) that emits explicit chains for debugging, planning, and agentic coding ### StepFun - [stepfun-ai/Step-3.7-Flash](https://recipes.vllm.ai/stepfun-ai/Step-3.7-Flash): Production-grade vision-language MoE (~198B total / 11B active parameters) combining a 196B sparse language backbone with a 1.8B perception encoder, hybrid SWA/Global attention, and 3-way Multi-Token Prediction - [stepfun-ai/Step-3.5-Flash](https://recipes.vllm.ai/stepfun-ai/Step-3.5-Flash): Production-grade reasoning MoE (~196B total / 11B active parameters) with hybrid attention schedules, SWA compensation, and multi-token prediction for low-latency long-context inference ### Jina AI - [jinaai/jina-embeddings-v5-text-small](https://recipes.vllm.ai/jinaai/jina-embeddings-v5-text-small): Jina AI's fifth-gen multilingual text embedding model (677M, Qwen3-0.6B-Base) with task-specific LoRA adapters for retrieval, text-matching, classification, and clustering. - [jinaai/jina-reranker-m0](https://recipes.vllm.ai/jinaai/jina-reranker-m0): Multilingual, multimodal reranker for text and visual documents across 29+ languages via Qwen2-VL backbone ### Wan (Alibaba) - [Wan-AI/Wan2.2-T2V-A14B-Diffusers](https://recipes.vllm.ai/Wan-AI/Wan2.2-T2V-A14B-Diffusers): Wan2.2 video generation models — T2V/I2V MoE (14B active) and unified TI2V (5B dense), served via vLLM-Omni ### Seed (ByteDance) - [ByteDance-Seed/Seed-OSS-36B-Instruct](https://recipes.vllm.ai/ByteDance-Seed/Seed-OSS-36B-Instruct): ByteDance Seed-OSS 36B dense model with unique 'thinking budget' control and 512K context support ### InternVL (OpenGVLab) - [OpenGVLab/InternVL3_5-8B](https://recipes.vllm.ai/OpenGVLab/InternVL3_5-8B): InternVL 3.5 vision-language models from Shanghai AI Lab with thinking-mode prompting ### Arcee AI - [arcee-ai/Trinity-Large-Thinking](https://recipes.vllm.ai/arcee-ai/Trinity-Large-Thinking): Arcee AI's reasoning-focused sparse MoE (AfmoeForCausalLM) with structured traces and agentic tool use ### LongCat (Meituan) - [meituan-longcat/LongCat-Image-Edit](https://recipes.vllm.ai/meituan-longcat/LongCat-Image-Edit): Bilingual (Chinese-English) image editing model from Meituan LongCat, served via vLLM-Omni ## Feeds - [Sitemap](https://recipes.vllm.ai/sitemap.xml): canonical XML sitemap. - [llms-full.txt](https://recipes.vllm.ai/llms-full.txt): full text of every recipe, concatenated for AI-assistant retrieval.