AuroraNemoia’s avatarAuroraNemoia’s Twitter Archive—№ 458

  1. llama.cpp server pushing 18.6t/s on a P40 from 2016, absolutely ridiculous, next you'll tell me a 33B model quantized would still yield good enough perf for my usecase.