Something I think would help accelerate this is educating people on how to run them. But maybe that comes after a long period of stability, since models leapfrog each other every few months now.
Like, first, which model do I run? Depending on your hardware, you’ll get 50 answers and random links to huggingface. Are the things they are linking models from companies or organizations you hear about? Maybe, or maybe they’re some strange distillation or uncensored/abliterated model, or some custom tweak that nobody can explain without using a bunch of nonsensical sounding terms.
Second, what are they actually capable of? What other apps can they link to and how do you link them? The second part is not always easy. A friend recently started using them to assist with organizing his life and thoughts by connecting it to Obsidian and Zotero. He has ADHD and says it helps him immensely in actually getting those ideas down, if not developing them further. I never would have considered this kind of workflow.
Third, what software do I use? LM Studio? Ollama? Ramalama? How come some connect to my IDE, but not others? What is rocm and why do some support it and some not? Do I need that?
These are rhetorical questions, but it’s honestly a swamp trying to get clear answers to them online and probably the biggest barrier to getting people to use these local models at the moment.
usually the merge models, uncensored ones, abliterated, etc are complete dogshit. if you can, run unsloth ggufs on llama.cpp and anything of vital importance in transformers
For sure, I think having an easy to follow process for people to get set up would be really useful. I can see how all of this can be pretty overwhelming, and unless you know what you can do with these things, it might not even feel like it’s worth the bother.
I expect in a year or so things will likely start to get a bit more stable, and the churn will slow down enough for some standard solutions to start emerging. Even if the models keep changing, having a standardized harness to drop them into would go a long way towards making this tech accessible.
Given that we’re nearly at the point where decade-old GPUs can run close to the best models, I think consumer-grade “LLM machines” are entirely plausible if hardware prices can ever stabilize. Could just be an eGPU and snazzy software stack if someone already has a laptop.
With how much “AI” companies want to charge an eGPU would already pay for itself in a few days for just one person at a ridiculous “AI-first” company.
Yeah now that hardware isn’t as much of an issue (I got a couple models working on my years-old android) my barrier is pretty much as you’ve described. I can’t think of a personal usecase since I don’t code much, and as such delving into all the complexities is a big ask as of right now.
So far the only utility I regularly get out of LLMs is searching for internet results, but I have no idea how to go about getting a local model to do so, or if it’s even possible with old hardware.
Spent like an hour messing around on an old Linux desktop with an rx580, but never figured it out lol
Something I think would help accelerate this is educating people on how to run them. But maybe that comes after a long period of stability, since models leapfrog each other every few months now.
Like, first, which model do I run? Depending on your hardware, you’ll get 50 answers and random links to huggingface. Are the things they are linking models from companies or organizations you hear about? Maybe, or maybe they’re some strange distillation or uncensored/abliterated model, or some custom tweak that nobody can explain without using a bunch of nonsensical sounding terms.
Second, what are they actually capable of? What other apps can they link to and how do you link them? The second part is not always easy. A friend recently started using them to assist with organizing his life and thoughts by connecting it to Obsidian and Zotero. He has ADHD and says it helps him immensely in actually getting those ideas down, if not developing them further. I never would have considered this kind of workflow.
Third, what software do I use? LM Studio? Ollama? Ramalama? How come some connect to my IDE, but not others? What is rocm and why do some support it and some not? Do I need that?
These are rhetorical questions, but it’s honestly a swamp trying to get clear answers to them online and probably the biggest barrier to getting people to use these local models at the moment.
usually the merge models, uncensored ones, abliterated, etc are complete dogshit. if you can, run unsloth ggufs on llama.cpp and anything of vital importance in transformers
For sure, I think having an easy to follow process for people to get set up would be really useful. I can see how all of this can be pretty overwhelming, and unless you know what you can do with these things, it might not even feel like it’s worth the bother.
I expect in a year or so things will likely start to get a bit more stable, and the churn will slow down enough for some standard solutions to start emerging. Even if the models keep changing, having a standardized harness to drop them into would go a long way towards making this tech accessible.
Given that we’re nearly at the point where decade-old GPUs can run close to the best models, I think consumer-grade “LLM machines” are entirely plausible if hardware prices can ever stabilize. Could just be an eGPU and snazzy software stack if someone already has a laptop.
With how much “AI” companies want to charge an eGPU would already pay for itself in a few days for just one person at a ridiculous “AI-first” company.
Yeah now that hardware isn’t as much of an issue (I got a couple models working on my years-old android) my barrier is pretty much as you’ve described. I can’t think of a personal usecase since I don’t code much, and as such delving into all the complexities is a big ask as of right now.
So far the only utility I regularly get out of LLMs is searching for internet results, but I have no idea how to go about getting a local model to do so, or if it’s even possible with old hardware.
Spent like an hour messing around on an old Linux desktop with an rx580, but never figured it out lol