Comments (4)
Should be fixed in the next release, by #4609.
from vllm.
Thanks, do you know when? or do you have an idea how I can go around this? It's happening when I run lora on a 70b model which is running on 2 GPU, I'm trying to load llama3 70b on a single GPU (a100) but it doesn't seem to work, I run our of vram...
from vllm.
The next release is just around the corner. If you can't wait for that, you can install vLLM from main
branch directly.
from vllm.
Fixed by #4609, which has been released in v0.4.3.
Edit: Technically it is still in pre-release but should be out very soon.
from vllm.
Related Issues (20)
- when i set tensor_parallel_size>1(A100 * 4), it does not work HOT 8
- [Bug]: `samplers/test_logprobs.py` fail on H100
- [Bug]: Timeout Error When Deploying Llamafied InternLM2-5-7B-Chat-1M Model via vLLM OpenAI API Server
- [Feature]: Apply chat template through `LLM` class HOT 11
- [Bug]: When using qwen-32b-chat-awq with multi-threaded access, errors occur after approximately several hundred visits.”vllm.engine.async_llm_engine.AsyncEngineDeadError: Background loop has errored already.“ HOT 1
- [Feature]: Return softmax of attention layer. HOT 3
- [Bug]: Paligemma support for PNG files HOT 5
- [Bug]: illegal memory access when increase max_model_length on FP8 models HOT 6
- [Bug]: autogen can't work with vllm v0.5.1
- v0.5.2, v0.5.3, v0.6.0 Release Tracker HOT 5
- [Bug]: Severe computation errors when batching request for microsoft/Phi-3-mini-128k-instruct HOT 3
- [Bug]: The shape of the embed_tokens of llama model doesn't match the llama3 configuration
- [Bug]: TypeError: 'NoneType' object is not callable when start Gemma2-27b-it HOT 5
- [Bug]: Seed issue with Pipeline Parallel HOT 2
- [Bug]: vLLM is unable to load Mistral on Inferentia and AWS neuron HOT 1
- [Installation]: ERROR: Could not find a version that satisfies the requirement pyzmq (from versions: none) HOT 4
- [Bug]: No metrics exposed at /metrics with 0.5.2 (0.5.1 is fine), possible regression? HOT 3
- [Bug]: Can't load gemma-2-9b-it with vllm 0.5.2 HOT 9
- unable to run vllm model deployment HOT 6
- [Bug]: failed when run Qwen2-54B-A14B-GPTQ-Int4(MOE) HOT 3
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from vllm.