Files
xSaherelmWorkspace/Documents/Ai Serving.txt
T
2026-10-03 18:04:36 +03:30

40 lines
1.3 KiB
Plaintext

download llamafile from github repo : https://github.com/mozilla-ai/llamafile/releases/tag/0.10.6
rename it to .exe
if using flash format it as extFAT
go to Hugging Face and downoad GGUF format of LLMS you want
open notepad and create a file named start.bat:
./llafial*.*.exe --server --model "llm-gguf-file.gguf" --ctx-size 32768
./llafial*.*.exe --server --embeddings --model "llm-gguf-file.gguf" --ctx-size 32768
--host 0.0.0.0
--port 8081
sc.exe create XLLMEmbdService binPath= "D:\Projects\LLMs\bge-m3.bat"
sc.exe create XLLMService binPath= "D:\Projects\LLMs\gemma3b1sc.bat"
sc queryex type=service state=all
sc queryex type=service state=all | find /i "SERVICE_NAME:"
sc queryex type=service state=active
sc query XLLMService
sc query XLLMEmbdService
sc delete XLLMService
sc delete XLLMEmbdService
https://docs.mozilla.ai/llamafile/using-llamafile/api
https://github.com/mozilla-ai/llamafile
-------------------
Uncesored Qwen 3.5 Coder GGUF:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
-------------------
Download these files and put them inside xAiApi/Models:
https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-base.bin
https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-small.bin