-
So, it is October 2nd 2023. What's the fastest way to run LLMs on CUDA? (Compute Capability 6.1, I know, old.) Transformers with Mistral-7B currently gets me 7.8t/s Limited to float16 or GPTQ. Though GPTQ seems to run noticeably slower, because it's emulating int4.
AuroraNemoia’s Twitter Archive—№ 361