DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale

vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale

Comments
5 min read
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches

vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches

Comments
4 min read
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Comments
4 min read
Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs

Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs

Comments
1 min read
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

15
Picked as gem Comments 1
9 min read
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

1
Comments
11 min read
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

1
Comments
8 min read
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

9
Comments
11 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Comments
9 min read
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

1
Comments
9 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Comments
13 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
Inside vLLM: Following One Request from the API to GPU Execution

Inside vLLM: Following One Request from the API to GPU Execution

1
Comments 2
24 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

7
Comments
9 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

13
Comments
13 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.