download llamafile from github repo : https://github.com/mozilla-ai/llamafile/releases/tag/0.10.6 rename it to .exe if using flash format it as extFAT go to Hugging Face and downoad GGUF format of LLMS you want open notepad and create a file named start.bat: ./llafial*.*.exe --server --model "llm-gguf-file.gguf" --ctx-size 32768 ./llafial*.*.exe --server --embeddings --model "llm-gguf-file.gguf" --ctx-size 32768 --host 0.0.0.0 --port 8081 sc.exe create XLLMEmbdService binPath= "D:\Projects\LLMs\bge-m3.bat" sc.exe create XLLMService binPath= "D:\Projects\LLMs\gemma3b1sc.bat" sc queryex type=service state=all sc queryex type=service state=all | find /i "SERVICE_NAME:" sc queryex type=service state=active sc query XLLMService sc query XLLMEmbdService sc delete XLLMService sc delete XLLMEmbdService https://docs.mozilla.ai/llamafile/using-llamafile/api https://github.com/mozilla-ai/llamafile ------------------- Uncesored Qwen 3.5 Coder GGUF: https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF ------------------- Download these files and put them inside xAiApi/Models: https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-base.bin https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-small.bin