AI 见闻
精选· 重要性 4/5

vLLM:面向LLM的高吞吐量、内存高效推理与服务引擎

GitHub Trending (AI repos)··vllm-project·约 1 分钟阅读
中文导读

vLLM是一个开源项目,为大型语言模型提供高吞吐量和内存高效的推理与服务引擎,显著提升部署效率。

vllm-project/vllm87,149

A high-throughput and memory-efficient inference and serving engine for LLMs

在 GitHub 查看完整介绍
原文出处
vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs

本文为机器翻译辅以 AI 润色,仅供参考。原始事实以原文为准。

相关阅读