Views
No views yet
--token-embedding-type & --output-tensor-type are in Q8_0) here:--reasoning off/-rea off should be added as llama.cpp cli arg to workaround an issue where the actual output is just in the reasoning part - I found that to be the case too and disabling reasoning helped.