Speculative Decoding in vLLM on AMD GPUs

(vllm.ai)

33 points | by ankitg12 2 hours ago

2 comments

  • intothemild 1 hour ago
    Whilst this is an excellent post from vLLM, one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.

    Stock vLLM runs so slowly on these cards compared with vLLM forks like Radiance. Going from say 20-30t/s gen, to 150-200t/s

    Most of AMD/vLLM work seems to be around their data centre cards, or the AMD AI Halo/Ryzen and ignores the R9700 AI Pro.

    Really wish this would change.

    • minraws 55 minutes ago
      I don't see a reason why it should AMD doesn't care about lower end prosumers atm.

      They might in the future but future is in the future ofc

      Edit: to be clear I think it's ridiculous they don't but from a company's stand point it doesn't make much sense

      • websap 48 minutes ago
        Yeah, as a business AMD should first care about getting their DC grade hardware optimized for inference workloads. It's unfortunate that most of HN discussion has devolved to me-ish.
    • dist-epoch 21 minutes ago
      George Hotz in June 2023:

      > I have had direct contact with members of the AMD RTG team and I was disgusted to find that AMD doesn't even provide them with hardware to work on. The developer I was working with had to buy the GPU he was writing drivers for.

      • da-x 11 minutes ago
        I think this has changed since then, their policies toward open source improved (e.g ROCm).
        • dist-epoch 5 minutes ago
          The market says the problem is still there.

          An NVIDIA consumer GPU sells for 50+% or more than an equivalent AMD GPU. Because people are buying NVIDIA GPUs to run local models instead of AMD ones.

          I did the same thing, I paid 50% more to get an 5070 Ti instead of the equivalent AMD.

          This is probably good for gamers, AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.

  • hn45e7pbij 1 minute ago
    [dead]