Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
Watermarking in vLLM (vllm.ai)
3 points by eatonphil 1 day ago | past | discuss
Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin (vllm.ai)
2 points by ankitg12 14 days ago | past
Speculative Decoding in vLLM on AMD GPUs (vllm.ai)
145 points by ankitg12 18 days ago | past | 53 comments
Efficient Decode Context Parallelism with vLLM for Long Context Workloads (vllm.ai)
1 point by aray07 27 days ago | past
vLLM Recipes (vllm.ai)
4 points by kristianpaul 52 days ago | past
Kimi K3 on vLLM: Up to 370 Tokens/sec (vllm.ai)
7 points by wskwon 60 days ago | past
vLLM prefill paired with TileRT decode (vllm.ai)
2 points by ilreb 73 days ago | past
Automatic Prefix Caching – vLLM (vllm.ai)
1 point by ankitg12 84 days ago | past
Micro-Agent: Beat Frontier Models with Collaboration Inside Model API (vllm.ai)
81 points by matt_d 88 days ago | past | 22 comments
vLLM Recipes (vllm.ai)
3 points by kristianpaul 3 months ago | past
Fast and Efficient LLM Inference with vLLM: A New Course with Deeplearning.ai (vllm.ai)
3 points by sonabinu 3 months ago | past
Session-Aware Agentic Routing: Continuity-Aware Model Selection for Long-Horizon (vllm.ai)
2 points by matt_d 3 months ago | past
Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team (vllm.ai)
69 points by berlianta 4 months ago | past | 24 comments
vLLM Semantic Router v0.2 Athena: ClawOS, Model Refresh, and the System Brain (vllm.ai)
3 points by mariuz 6 months ago | past
GPT-OSS Optimizations on Nvidia Blackwell: Pushing the Pareto Frontier (vllm.ai)
2 points by roody_wurlitzer 7 months ago | past
vLLM WideEP and Large-Scale Serving Toward Maturity on Blackwell (Part I) (vllm.ai)
1 point by roody_wurlitzer 7 months ago | past
DeepSeek-v3.2 on GB300: Performance Breakthrough (vllm.ai)
2 points by roody_wurlitzer 7 months ago | past
vLLM large scale serving: DeepSeek 2.2k tok/s/h200 with wide-ep (vllm.ai)
147 points by robertnishihara 8 months ago | past | 54 comments
VLLM: The High-Throughput and Memory-Efficient Serving Engine for LLMs (vllm.ai)
1 point by sorrow17 9 months ago | past
Bitwise Consistent On-Policy Reinforcement Learning with VLLM and TorchTitan (vllm.ai)
1 point by brrrrrm 10 months ago | past
vLLM TPU: A New Unified Backend Supporting PyTorch and JAX on TPU (vllm.ai)
1 point by pykello 11 months ago | past
VLLM TPU: A New Unified Back End Supporting PyTorch and Jax on TPU (vllm.ai)
1 point by alphabetting 11 months ago | past
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (vllm.ai)
2 points by matt_d on Sept 12, 2025 | past
vLLM with torch.compile: Efficient LLM inference on PyTorch (vllm.ai)
1 point by matt_d on Sept 4, 2025 | past
VLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention (vllm.ai)
20 points by jxmorris12 on July 2, 2025 | past | 5 comments
VLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention (2023) (vllm.ai)
3 points by telotortium on March 29, 2025 | past
vLLM V1: A Major Upgrade to vLLM's Core Architecture (vllm.ai)
2 points by ozgune on Jan 31, 2025 | past
vLLM V1: A Major Upgrade to vLLM's Core Architecture (vllm.ai)
5 points by xmo on Jan 27, 2025 | past
VLLM 2024 Retrospective and 2025 Vision (vllm.ai)
1 point by shenli3514 on Jan 19, 2025 | past
Installing and Developing VLLM with Ease (vllm.ai)
1 point by brethil on Jan 13, 2025 | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: