Views
No views yet
cat command.launch.py script from Distributed Llama repository: python launch.py llama3_1_405b_instruct_q40make dllama./dllama chat --model dllama_model_llama31_405b_q40.m --tokenizer dllama_tokenizer_llama_3_1.t --buffer-float-type q80 --max-seq-len 2048 --nthreads 64