IT
Accedi Pubblica
contextetech /

llama.cpp server for a GGUF model

v1
Inglese Licenza: MIT Pubblicato il aggiornato 1 ora fa 0 utilizzi

Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port 8080, works without a GPU too.

Modellomodel.gguf
Hardware testatoCPU only, or a GPU (use -ngl to offload layers)

Comando di avvio

llama-server -m model.gguf -c 8192 --host 127.0.0.1 --port 8080

Community

Ancora nessun commento. Condividi la tua esperienza, aiuterà chi viene dopo.