Okay akutally Qwen 3.8 might be the first qwen model that I can actually work with. All new models have a bias but i've been using it for a few hours so far today. A little slow, but hopefully I can get that tweaked. I lose cloud tokens at the end of next week, so lets see if it can keep up :)
27b-mtp-q8_0
kv: 256k-q8_0 - 41gb in use.
Login to reply
Replies (2)
You should allegedly be able to run it with 1M context window with yarn. Havnt tried yet myself.
Cloud tokens expire, control remains, courtesy of Bezos and pals.