6 ms·
ChatGPT-J: The Privacy-First, Self-Hosted Chatbot Built on GPT-J's Powerful AI
- two_handfuls 4y ago"Privacy-First", but also working in a colab notebook - meaning running on someone else's machine? That doesn't seem very private.
- was_a_dev 4y agoDownload the notebook and run locally?
- jarrell_mark 4y agoYes, the GitHub has the Jupyter .ipynb notebook that can be run locally: https://github.com/jarrellmark/chatgpt-j https://github.com/jarrellmark/chatgpt-j And even in Colab, it's privacy first in the sense that user input or model output isn't being sent anywhere. The data is local to your Colab session.
- johntash 4y agoCan this be run locally without beefy GPUs by any chance?
- jarrell_mark 4y agoIt should work with about 12gb GPU RAM. I got it to load on a GTX 1070 with 8GB GPU RAM, but then it crashed before it could generate a response. It needs less RAM than regular GPT-J because the weights are converted to 8-bit
- ops 4y agoI haven't used this yet, but I am currently running GPT-J on my Mac Studio, so I suspect so.
- jerpint 4y agoThere have been CPU implementations of LLAMA (7b parameters, comparable in size) with very impressive performance
- quesomaster9000 4y agoggml (https://github.com/ggerganov/ggml https://github.com/ggerganov/ggml) has a GPT-J example, the 6B parameter model runs happily on the CPU 16gb of ram and 8 cores at a couple of words per second, no GPUs necessary. gptj_model_load: ggml ctx size = 13334.86 MB gptj_model_load: memory_size = 1792.00 MB, n_mem = 57344 gptj_model_load: model size = 11542.79 MB / num tensors = 285 main: number of tokens in prompt = 12 An example of GPT-J running on the CPU is shown in Fig. [4](#Fig4 main: mem per token = 16179460 bytes main: load time = 7463.20 ms main: sample time = 3.24 ms main: predict time = 4887.26 ms / 232.73 ms per token main: total time = 13203.91 ms
- serendipty01 4y agoGithub repo : https://github.com/jarrellmark/chatgpt-j https://github.com/jarrellmark/chatgpt-j