No, they’re hoarding them so you have to pay for cloud services they control from now on. With your little Fire tablet. No more pirating movies or political organizing for you, piggy
9060xt 16gb is the most cost effective new GPU, but if you’re going used look for a V620 on eBay. It’s a 6800xt chip but in server form factor GPU with 32GB vram. Can be a bit of a pain to set up but by far the most cost effective option IMO
Local 27b models are good enough for most tasks.
Can’t wait to buy one of these from Ebay for 10% of the price next year.
Yeah no way they will allow any of this hardware to go back onto the market. Anything they dont use anymore will be destroyed.
Buy it, destroy it. Just like buying bunch of old books, train their LLM’s and burn it. Humanity has gone a long way to be that stupid.
https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.
Upgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.
I plan to once there is a version with turboquant and MTP as that huge context window is key.
And that’s really why they’re hoarding them.
No, they’re hoarding them so you have to pay for cloud services they control from now on. With your little Fire tablet. No more pirating movies or political organizing for you, piggy
For what it’s worth, piracy is the primary thing I use my older 10" fire tablet for.
Any 27b Model you can currently recommend for a 16gb AMD ? Mostly coding tasks but not exclusively.
9060xt 16gb is the most cost effective new GPU, but if you’re going used look for a V620 on eBay. It’s a 6800xt chip but in server form factor GPU with 32GB vram. Can be a bit of a pain to set up but by far the most cost effective option IMO
V620s were a good deal when you could get them for $350, now they’re $700+ and no longer a good deal.
For your hardware, the VRam is not enough to run 27b but, I’d recommend Qwen 3.5 9b for image / text to text.
And I’m planning to experiment with Qwen 3.8 9b for text to text.
4_k_m quantization is the sweet spot for performance and ram usage.
Also, I find Llama cpp is better than Ollama in terms of performance.