5 ms·
My M5 Pro gets around 12-15 (6 bit MTP), although I haven’t worked on optimising it at all yet. A nice thing about running locally is you can run an uncensored
by trollbridge 25d ago
My M5 Pro gets around 12-15 (6 bit MTP), although I haven’t worked on optimising it at all yet.
A nice thing about running locally is you can run an uncensored model and you don’t have to worry about TOS violations on your OpenAI account when you ask it to “reverse engineer this ancient router firmware and give me a licence key that will work on it”.
- medler 25d agoQwen is very much censored. Just try asking it about Tiananmen or how to build a bomb. But it is nice that you can experiment with it locally without having to worry about your account getting nuked
- jchw 25d agoYou are misunderstanding what they said, they are saying you can use uncensored variants of models like Qwen when running locally. There are quite a lot of people working to "uncensor" open weights releases. It seems to work although it would be nice if some third party was benchmarking the uncensored variants regularly to give us an idea of how well retained their skills are.
- trollbridge 24d agoAbliteratuon sloghtly reduces the strength of the model - in my opinion it’s around the same jump as going from 5 bit to 4 bit.