Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
c0rruptbytes
1 day ago
|
parent
|
context
|
favorite
| on:
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB ...
so many inference project, omlx already supports all of this and has a 1000 people trying to optimize it constantly
help
carloslfu
1 day ago
|
next
[–]
Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX
reply
carloslfu
1 day ago
|
prev
[–]
Interesting! I'll check it out
reply
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: