Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
nullc
3 months ago
|
parent
|
context
|
favorite
| on:
Accelerating Gemma 4: faster inference with multi-...
Thanks for the link,it took qwen3.6-27B-q8 w/256k context on my RTX A6000 from ~20t/s to 55t/s. Prefill is mysteriously slower however, but prefill is so much faster still that I think I'm still bottlenecked on output most of the time.
_factor
3 months ago
[–]
Took 2x AMD MI50s to 50 t/s instead of 20 t/s for Q8 27B. Impressive.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: