Rendered at 09:55:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
antonyragleap 26 minutes ago [-]
Curious about Vulkan overhead on Intel vs AMD/Nvidia for long context. Any benchmarks vs vllm/sglang?
PcChip 13 hours ago [-]
I didn't see any benchmarks against vllm, sglang, exllama, etc
rancor 12 hours ago [-]
Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.
nullpoint420 6 hours ago [-]
Woof. Wonder if the creator knows that
dlcarrier 11 hours ago [-]
From what I've seen, Vulkan adds a lot of overhead on Intel hardware.
gunalx 3 hours ago [-]
Yes and no. I tested llamacpp on my intel gpu with both vulkan and sycl. I measured them to be in the same ballpark even if i have se en pepole claim marger differences than i observed.
In the end i took the sligth slowdown of vulkan to have a more stable and higher development velocity backend
While being able to use identical setups on both intel and and gpus.
aidiveyt 3 hours ago [-]
claude code appends a role:"system" block after the user prompt, so a proxy rewriting the trailing user message is a no-op.
DylanMerigaud 3 hours ago [-]
Intel hardware can have Vulkan overhead, impacting performance.
peddling-brink 11 hours ago [-]
> llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
I got excited about someone paying attention to intel. Oh well.
wronglebowski 8 hours ago [-]
What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.
peddling-brink 8 hours ago [-]
Two arc b60s. The intel vllm build is getting me ~15t/s decode with heavy context using qwen3.8 27b.
kamranjon 10 hours ago [-]
llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags
peddling-brink 7 hours ago [-]
Llama would be nice for the ggufs. Any specific flags or tutorials I should look at?
In the end i took the sligth slowdown of vulkan to have a more stable and higher development velocity backend While being able to use identical setups on both intel and and gpus.
I got excited about someone paying attention to intel. Oh well.
https://github.com/ggml-org/llama.cpp/blob/master/docs/backe...