• deleted@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    13 天前

    Local 27b models are good enough for most tasks.

    Can’t wait to buy one of these from Ebay for 10% of the price next year.

      • berty@feddit.org
        link
        fedilink
        English
        arrow-up
        2
        ·
        12 天前

        Buy it, destroy it. Just like buying bunch of old books, train their LLM’s and burn it. Humanity has gone a long way to be that stupid.

        • Lydia_K@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          arrow-down
          1
          ·
          12 天前

          I plan to once there is a version with turboquant and MTP as that huge context window is key.

      • 4am@lemmy.zip
        link
        fedilink
        English
        arrow-up
        3
        ·
        12 天前

        No, they’re hoarding them so you have to pay for cloud services they control from now on. With your little Fire tablet. No more pirating movies or political organizing for you, piggy

    • Chee_Koala@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      edit-2
      12 天前

      Any 27b Model you can currently recommend for a 16gb AMD ? Mostly coding tasks but not exclusively.

      • mierdabird@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        0
        ·
        12 天前

        9060xt 16gb is the most cost effective new GPU, but if you’re going used look for a V620 on eBay. It’s a 6800xt chip but in server form factor GPU with 32GB vram. Can be a bit of a pain to set up but by far the most cost effective option IMO

        • Darkaga@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          12 天前

          V620s were a good deal when you could get them for $350, now they’re $700+ and no longer a good deal.

      • deleted@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        12 天前

        For your hardware, the VRam is not enough to run 27b but, I’d recommend Qwen 3.5 9b for image / text to text.

        And I’m planning to experiment with Qwen 3.8 9b for text to text.

        4_k_m quantization is the sweet spot for performance and ram usage.

        Also, I find Llama cpp is better than Ollama in terms of performance.