# User guides How-to guides for deploying, scaling, and operating Ray Serve LLM. If you are new, start with the {doc}`Quickstart <../quick-start>`, then come back here to go deeper. ## Configure and deploy - {doc}`Configuration reference `: every `LLMConfig` field, from model loading and engine kwargs to accelerators, placement, and deployment options. - {doc}`Deployment initialization `: speed up model loading and replica startup with caching, streaming load formats, and initialization callbacks. - {doc}`Multi-LoRA deployment `: serve many LoRA adapters on a shared base model with runtime switching and an LRU cache. ## Scale across GPUs and nodes - {doc}`Cross-node parallelism `: distribute a model across GPUs and nodes with tensor and pipeline parallelism and placement groups. - {doc}`Data parallel attention `: replicate the model into coordinated data-parallel groups to raise throughput, especially for MoE models. - {doc}`Fractional GPU serving `: pack multiple small-model replicas onto a single GPU. ## Optimize latency and throughput - {doc}`Prefill/decode disaggregation `: split prompt processing and token generation onto separate replicas to tune each independently. - {doc}`KV cache offloading `: extend KV cache capacity with LMCache and tiered storage backends. - {doc}`Prefix-aware routing `: route requests to replicas that already hold a matching prefix to maximize cache hits. - {doc}`Direct streaming `: bypass the ingress when streaming tokens to cut per-token latency. ## Choose an engine - {doc}`vLLM compatibility `: use vLLM features such as embeddings, structured outputs, vision, and reasoning through Ray Serve LLM. - {doc}`Custom vLLM models `: serve an out-of-tree architecture with a vLLM plugin, using a Qwen3 reward model as the example. - {doc}`SGLang integration `: run SGLang as the inference engine instead of vLLM. ## Operate in production - {doc}`Observability and monitoring `: engine and request metrics, Grafana dashboards, and Prometheus integration. ```{toctree} :hidden: :maxdepth: 1 Configuration reference Deployment initialization Multi-LoRA deployment Cross-node parallelism Data parallel attention Fractional GPU serving Prefill/decode disaggregation KV cache offloading Prefix-aware routing Direct streaming vLLM compatibility Custom vLLM models SGLang integration Observability and monitoring ```