DE
Anmelden Veröffentlichen
contextetech /

Qwen2.5 7B with vLLM on a 24 GB GPU

v1
Englisch Lizenz: MIT Veröffentlicht am aktualisiert vor 1 Stunde 0 Nutzungen

Serve Qwen2.5 7B Instruct with vLLM on a 24 GB GPU: OpenAI-compatible API, 8k context, a single command.

ModellQwen/Qwen2.5-7B-Instruct
Getestete HardwareNVIDIA GPU with 24 GB (for example RTX 4090 or L4)

Startbefehl

vllm serve Qwen/Qwen2.5-7B-Instruct --max-model-len 8192 --gpu-memory-utilization 0.90 --port 8000

Community

Noch keine Kommentare. Teile deine Erfahrung, sie hilft den Nächsten.