6 ms·
Hi, I am the creator, feel free to ask any questions :) What do you think about it?
by gioscarab 12d ago
Hi, I am the creator, feel free to ask any questions :)
What do you think about it?
- mpalmer 12d agoAt first blush, it is a really persuasive compromise between full-on LLM inference and boring old fuzzy history search! I really like it, this flavor of specialization gives the user a win on privacy and speed. Seems like the right idea for such a tool.
- registereduser1 12d agoCool project! How does it differ from warp terminals ai mode where you can ask it questions and it responds back
- gioscarab 12d agoWarp uses LLMs so it is slow and prone to hallucination. Using very colloquial terms TERMy is more or less a calculator that knows english :) so it can run on your CPU and respond instantly! The difference is that it can only answer predetermined responses (with optional arguments) this makes it useless if you need to generate text, but makes it safe and predictable for a use case like a terminal assistant.
- gurjeet 12d agoI haven't evaluated it yet, but I love the fact that the output is (at least claimed to be) deterministic. I can't trust an LLM to do the right thing after I deploy it to production, because their output is non-deterministic by design. TERMy (or is it the NPC-forge) seems to be worth a try.
- kouteiheika 12d ago> because their output is non-deterministic by design. It isn't. At least not by design, even though in practice it often can be. If you do greedy decoding (or use a preset seed) and deterministically compute everything (e.g. only use integer math) then it will be 100% always deterministic.
- kennywinker 12d agoThat’s true, but not true-true. Sure, every time you prompt “what is the weather in kansas” you’ll get the same output, but if you prompt “what is the weather in kansas right now” you’ll get a different output, and then “what is the weather in kansas today” gets a different output. Language being language, there are infinite ways to say things, so there are infinite variations in what the llm can output in response to very similar prompts. This tool has a finite amount of outputs for an infinite amount of inputs. Which is different from an llm based tool.
- skeledrew 12d agoI think the point being made is that given a particular input string, you can get a deterministic output string back from the LLM.
- kennywinker 12d agoYes i think I acknowledged that, but is that useful for making a tool that can be trusted to safely run shell commands when asked arbitrary questions? No. It’s not.
- piterrro 12d agoYou can get determinostic output (mostly) by setting the temperature to zero. Using couple of other tricks you can get close to 100% of determinism with LLMs.
- jdiff 12d agoThat's reproducible, I wouldn't call it deterministic. Small, semantically meaningless changes in the input can still result in wildly different output.
- kzrdude 12d agoI've long observed that kind of behaviour in google translate (which makes sense, they have been using ML for a long time.)
- asQuirreL 12d agoThat's the definition of a chaotic system (small change in initial conditions results in large, seemingly -- but not actually -- random changes in output), but it's still deterministic (same input results in same output).
- utopiah 12d agoWhat dataset does step 5 rely on? Is it from your own terminal history, man pages, scrapped dataset from e.g. StackOverflow, sth else?
- gioscarab 12d agoThe dataset is here: https://github.com/gioblu/NPC-Forge/tree/main/npcs/termy/dataset https://github.com/gioblu/NPC-Forge/tree/main/npcs/termy/dat... I hope the community will help me to enhance it :) it is just a proof of concept for now
- deleted 12d ago[deleted]
- kouteiheika 12d ago> Models like ornith:9b, mistral:7b or cogito:14b can get the job done sometimes, but they are not fast and reliable enough for general use, specially if you have only 4GB of VRAM. Have you considered/tried using a model that's, well, more appropriate size-wise for an use case like this? These are relatively big. Something like FunctionGemma [1] finetuned for a given set of tasks would be a lot more speedy. [1] https://blog.google/innovation-and-ai/technology/developers-tools/functiongemma/ https://blog.google/innovation-and-ai/technology/developers-...
- gioscarab 12d agoI tried functiongemma, it is for sure faster than those models, the problem is that is not reliable enough for a terminal assistant. I would say that no LLM is good for a terminal assistant, if you take into account the operational cost and the risk of damage. Even if it fails only 1 time out of 10 becomes useless. That's why I developed FlintParser!
- coder543 12d agoFunctionGemma never worked well for me (without fine tuning). Liquid has released 230M and 350M models that work far, far better in my testing: https://huggingface.co/LiquidAI/LFM2.5-230M https://huggingface.co/LiquidAI/LFM2.5-230M I really look forward to a hypothetical LFM3-230M, because LFM2.5-230M is so close to being usable, while FunctionGemma is miles away from being usable. But, yes, still tangential to TERMy.
- kennywinker 12d agohttps://github.com/ThorOdinson246/whatisit-nl2sh https://github.com/ThorOdinson246/whatisit-nl2sh uses a finetune of Qwen2.5-Coder-1.5B-Instruct. It works pretty well, tho it will misunderstand things from time to time
- tgv 12d agoCool, but system and user should probably stick to short, clear commands. E.g., I see you do some anaphora resolution (in particular: find what "it" refers to), but in a complex dialog, the human intention can differ from the machine's understanding. That will give problems when you end your dialog with "delete it". Adding more sentences to your data set will slowly degrade performance. It's a delicate system. Source: I have written software with similar functionality (NLP search) in SaaS form, a long time ago. It required quite a bit of work to configure.
- lna_stub 11d ago[flagged]
- cyberclimb 11d agois it supported to have it propose a command for approval rather than running autmatically? in the YT video it looks likw it ran the cpu temp command on its own love the idea/simplicity of this tool!
- gioscarab 11d agoYep he runs on its own when the command is non-destructive, like checking the CPU temperature, it does ask for permission if the command is potentially destructive. I agree it is so cool, it looks sci-fi :)