40 lines
1.3 KiB
Plaintext
40 lines
1.3 KiB
Plaintext
|
|
download llamafile from github repo : https://github.com/mozilla-ai/llamafile/releases/tag/0.10.6
|
|
rename it to .exe
|
|
if using flash format it as extFAT
|
|
go to Hugging Face and downoad GGUF format of LLMS you want
|
|
open notepad and create a file named start.bat:
|
|
./llafial*.*.exe --server --model "llm-gguf-file.gguf" --ctx-size 32768
|
|
./llafial*.*.exe --server --embeddings --model "llm-gguf-file.gguf" --ctx-size 32768
|
|
|
|
--host 0.0.0.0
|
|
--port 8081
|
|
|
|
sc.exe create XLLMEmbdService binPath= "D:\Projects\LLMs\bge-m3.bat"
|
|
sc.exe create XLLMService binPath= "D:\Projects\LLMs\gemma3b1sc.bat"
|
|
|
|
sc queryex type=service state=all
|
|
sc queryex type=service state=all | find /i "SERVICE_NAME:"
|
|
sc queryex type=service state=active
|
|
|
|
sc query XLLMService
|
|
sc query XLLMEmbdService
|
|
|
|
sc delete XLLMService
|
|
sc delete XLLMEmbdService
|
|
|
|
https://docs.mozilla.ai/llamafile/using-llamafile/api
|
|
https://github.com/mozilla-ai/llamafile
|
|
|
|
-------------------
|
|
|
|
Uncesored Qwen 3.5 Coder GGUF:
|
|
|
|
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
|
|
|
|
-------------------
|
|
|
|
Download these files and put them inside xAiApi/Models:
|
|
|
|
https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-base.bin
|
|
https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-small.bin |