59 ms·
Code Llama, a state-of-the-art large language model for coding
- mymac 3y agoNever before in the history of mankind was a group so absolutely besotted with the idea of putting themselves out of a job.
- ttul 3y agoThat’s just one perspective… Another perspective is that LLMs enable programmers to skip a lot of the routine and boring aspects of coding - looking up stuff, essentially - so they can focus on the fun parts that engage creativity.
- mymac 3y agoBut it won't stop there. Why would it stop at some arbitrarily defined boundary? The savings associated with no longer having to pay programmers the amounts of money that they believe they are worth (high enough to result in collusion between employers) are just too tempting.
- lsmeducation 3y agoOkay, think about it this way. This thing helps generate tons and tons of code. The more code people (or this thing) writes, the more shit there is to debug. More and more code, each calling each other means more and more insane bugs. We’re going to move from debugging some crap the last developer wrote to debugging an order of magnitude more code the last developer generated. It’s going to be wonderful for job prospects really.
- Sakthimm 3y agoUntil AI figures out debugging
- ilaksh 3y agoSome form of AI will eventually take over almost all existing jobs. Whether those jobs evolve or not somehow and new jobs replace them, we will see. But it's definitely not just programmers. And it will take time. Society needs to adjust. Stopping progress would not be a solution and is not possible. However, hopefully we can pause before we create digital animals with hyperspeed reasoning and typical animal instincts like self-preservation. Researchers like LeCun are already moving on from things like LLMs and working on approaches that really imitate animal cognition (like humans) and will eventually blow all existing techniques out of the water. The path that we are on seems to make humans obsolete within three generations or so. So the long term concern is not jobs, but for humans to lose control of the planet in less than a century. On the way there we might be able to manage a new golden age -- a crescendo for human civilization.
- lsmeducation 3y agoContinuing your aside… Humans don’t become obsolete, we become bored. This tech will make us bored. When humans get too bored and need shit to stir up, we’ll start a war. Take US and China, global prosperity is not enough right? We need to stoke the flames of war over Taiwan. In the next 300 years we’ll wipe out most of each other in some ridiculous war, and then rebuild.
- ilaksh 3y agoI agree that WWIII is a concern but I don't think it will be brought about by boredom. "Global prosperity" might be true in a very long-term historical sense, but it's misleading to apply it to the immediate situation. Taiwan is not just a talking point. Control over Taiwan is critical for maintaining hegemony. When that is no longer assured, there will likely be a bloody battle before China is given the free reign that it desires. WWIII is likely to fully break out within the next 3-30 years. We don't really have the facilities to imagine what 300 years from now will look like, but it will likely be posthuman.
- lsmeducation 3y agoI’ll go with the 30 year mark. Countries like Russia or China don’t get humbled in a loss (like Germany didn’t in WW1). Russia will negotiate some terms for Ukraine (or maintain perpetual war), but I believe it will become a military state that will funnel all money into the defense sector. The same with Iran, and the same with China. Iran supplies Russia with drones. I can promise you Russia will help Iran enrich their uranium. They are both pariah states, what do they have to lose? Nuclear Iran, here enters Israel. Everyone’s arming up, there’s a gun fight coming.
- swader999 3y agoThe answer to AI stealing your job is to go ahead and start a company, solve a hard problem, sell the solution and leverage AI to do this.
- astrange 3y agoThe only thing that takes anyone's job is demand shortfalls. Productivity increases certainly don't do it. It's like saying getting a raise makes you poorer.
- xcf_seetan 3y agoActually, in my country, Portugal, if your salary is the Minimum wage, you are exempt from paying taxes, but if you get a raise as little as 15€ or so, you move one category up and start to pay taxes, and you will receive less money than what you used to get before the raise.
- KingOfCoders 3y agoOne coachman to the other: "Another perspective about this car thing, you can skip all the routine and boring trips - they are done with cars. You can focus on the nice trips that make you feel good".
- yborg 3y agoWhen mechanized textile machinery was invented, the weavers that had jobs after their introduction were those that learned how to use them.
- thewataccount 3y agoIs automation not what every engineer strives for when possible? Especially software developers. From my experience with github copilot and GPT4 - developers are NOT going anywhere anytime soon. You'll certainly be faster though.
- kbrannigan 3y agoDo you really want to spend you days writing REDUX accumulators?
- 037 3y agoI understand the fear of losing your job or becoming less relevant, but many of us love this work because we're passionate about technology, programming, science, and the whole world of possibilities that this makes... possible. That's why we're so excited to see these extraordinary advances that I personally didn't think I'd see in my lifetime. The fear is legitimate and I respect the opinions of those who oppose these advances because they have children to provide for and have worked a lifetime to get where they are. But at least in my case, the curiosity and excitement to see what will happen is far greater than my little personal garden. Damn, we are living what we used to read in the most entertaining sci-fi literature! (And that's not to say that I don't see the risks in all of this... in fact, I think there will be consequences far more serious than just "losing a job," but I could be wrong)
- PUSH_AX 3y agoWe're not looking at a product that's putting anyone out of a job though, we're looking at a product that frees up a lot of time, and time is great.
- worksonmine 3y agoThis should be the only goal of mankind so we can smell the flowers instead of wasting our years in some cubicle. Some people will always want to work, but it shouldn't be the norm. What's the point really unless we're doing something we're passionate about? The economy?
- quickthrower2 3y agoThe best interpretation of this is you mean eventually ML/AI will put programmers out of a job, and not Code LLama specifically. However it is hard to tell how that might pan out. Can such an ML/AI do all the parts of the job effectively? A lot of non-coding skill bleed into the coder's job. For example talking to people who need an input to the task and finding out what they are really asking for, and beyond that, what the best solution is that solves the underlying problem of what they ask for, while meeting nonfunctional requirements such as performance, reliability, code complexity, and is a good fit for the business. On the other hand eventually the end users of a lot of services might be bots. You are more likely to have a pricing.json than a pricing.html page, and bots discover the services they need from searches, negotiate deals, read contracts and sue each other etc. Once the programming job (which is really a "technical problem solver" job) is replaced either it will just be same-but-different (like how most programmers use high level languages not C) or we have invented AGI that will take many other jobs. In which case the "job" aspect of it is almost moot. Since we will be living in post-scarcity and you would need to figure out the "power" aspect and what it means to even be sentient/human.
- vunderba 3y agoIf we get to the point where these large language models can create complete applications and software solutions from design specs alone, then there's no reason to believe that this would be limited to merely replacing software devs. It would likely impact a far larger swath of the engineering / design industry.
- astrange 3y agoYou can't get promoted unless you put yourself out of a job.
- drcongo 3y agoWell, since Brexit anyway.
- gw67 3y agoIn your opinion, Why Meta does this?
- chaorace 3y agoTo a certain extent, I think it's just IBM disease. A company the size of Meta is expected to have an AI research department like Microsoft or Google, even if their core business (social media) derives relatively less benefit from the technology. Pretend you're an uncreative PM on an AI team; what part of Facebook or VR could you feasibly improve by iterating on LLMs? Perhaps the content moderation system... but that would require wrangling with the company ethics comittee and someone else at the company probably already took ownership that idea. You've gotta do something compelling or else your ML engineers are going to run off somewhere else. If I were to ask my ML engineers about what they wanted to work on, they're going to avoid areas where their model is outgunned (i.e.: chat) and instead prefer lower hanging fruit which generalizes well on a resume (i.e.: "Pioneered and published key innovations in LLM code-generation"). Of course, the alternative answer is that Meta wants to replace all of their jr. developers with GPUs, but I think their leadership is a little too preoccupied with VR to be actively pushing for such a transformative initiative in anything more than a very uninvested capacity (e.g.: "Sure I'll greenlight this. Even if it doesn't pay off I don't have any better ideas")
- waitingkuo 3y agoLooks like that we need to request the access first
- taylorbuley 3y agoIn the past, LLAMA access was granted nearly immediately. For HuggingFace downloads, it took a full day.
- Havoc 3y agoThe bloke on hugging face usually has quantised versions minus the legal form
- reacharavindh 3y agoCode llama Python is very interesting. Specifically tuned for Python. I wonder if we could make such specific LLMs (one that is proficient in all things Rust, another- all things Linux, all things genomics, all things physics modeling etc) and have them talk to each other to collaboratively solve problems. That would be a crazy future thing! Putting machines truly to work..
- brucethemoose2 3y agoIf you can find a large body of good, permissively licensed example code, you can finetune an LLM on it! There was a similar attempt for Godot script trained a few months ago, and its reportedly pretty good: https://github.com/minosvasilias/godot-dodo https://github.com/minosvasilias/godot-dodo I think more attempts havent been made because base llama is not that great at coding in general, relative to its other strengths, and stuff like Starcoder has flown under the radar.
- deleted 3y ago[deleted]
- esperent 3y agoI think this is called "mixture of experts" and also there's a lot of speculation that it's how GPT-4 works, although probably with just a few large models rather than many small ones.
- jmiskovic 3y agoIt's been confirmed by multiple (unofficial) sources that GPT-4 is 8 models, each 220B parameters. Another rumor is GPT-4 being 16x111B models. There's a quite fresh and active project replicating something similar with herd of llamas: https://github.com/jondurbin/airoboros https://github.com/jondurbin/airoboros
- codercowmoo 3y agoCan you provide some of these sources?
- msoad 3y agoIs there any place we can try those models? Are they available on HuggingFace?
- maccam912 3y agoIt appears we do have a 34B version now, which never appeared for non fine tuned llama 2.
- jspisak 3y agoIt would be interesting to understand if a ~30B Llama-2 model would be interesting and for what reasons.
- hnuser123456 3y agoIt would fit on the 24GB top-end consumer graphics cards with quantization.
- brucethemoose2 3y agoLlama 34B is just big enough to fit on a 24GB consumer (or affordable server) GPU. Its also just the right size for llama.cpp inference on machines with 32GB RAM, or 16GB RAM with a 8GB+ GPU. Basically its the most desirable size for AI finetuning hobbyists, and the quality jump from llama v1 13B to llama v1 33B is huge.
- Tostino 3y agoBetter reasoning and general performance than 13b by far (if llama1 was any indication), and like the other user said, can fit on a single 24gb vram gaming card, and can be peft fine-tuned with 2x 24gb cards.
- airgapstopgap 3y agoLlama-1-33B was trained on 40% more tokens than LLama-1-13B; this explained some of the disparity. This time around they both have the same data scale (2T pretraining + 500B code finetune), but 34B is also using GQA which is slightly more noisy than MHA. Furthermore, there have been some weird indications in the original LLama-2 paper that 34B base model is something… even more special, it's been trained on a separate internal cluster with undervolted/underclocked GPUs (though this in itself can't hurt training results), its scores are below expectations, it's been less "aligned". Here, Code-Llama-Instruct-13B is superior to 34B on HumanEval@1. So yes, it's desirable but I wouldn't get my hopes up.
- lolinder 3y agoDoes anyone have a good explanation for Meta's strategy with AI? The only thing I've been able to think is they're trying to commoditize this new category before Microsoft and Google can lock it in, but where to from there? Is it just to block the others from a new revenue source, or do they have a longer game they're playing?
- deleted 3y ago[deleted]
- jspisak 3y agoIf you watch the Connect talks, I'll be speaking about this..
- lolinder 3y agoSorry—who are you and what are the Connect talks? I haven't heard of them and you don't have a bio.
- titaniumtown 3y agowhat is that?
- darrenf 3y agoFacebook Connect is what used to be called Oculus Connect. Kinda their equivalent of Apple's WWDC, I guess. It's when and where the Quest 3 will be officially unveiled in full, for example.
- modeless 3y agoInteresting that there's a 34B model. That was missing from the original Llama 2 release. I wonder if it's still usable for general non-code chat tasks or if the code fine tuning destroyed that. It should be the best model that would still fit on 24GB gaming GPUs with quantization, because 70B doesn't fit.
- brucethemoose2 3y agoSomeone "grafted" llama 33B onto llama v2 13B to make "llama 22B" https://huggingface.co/chargoddard/llama2-22b https://huggingface.co/chargoddard/llama2-22b Theoretically this is an even better size, as it would fit on a 20GB-24GB GPU with more relaxed quantization and much longer context. Metrics are slightly below 13B, but the theory is that the higher parameter count is more amenable to finetuning. If you search for 22B on huggingface, you can see that frankenllama experiments are ongoing: https://huggingface.co/models?sort=modified&search=22b https://huggingface.co/models?sort=modified&search=22b
- redox99 3y agoI can't imagine it being better than Llama1 33B, after all this code finetuning.
- modeless 3y agoBut the license for llama 2 is a whole lot better.
- redox99 3y agoMeh. If you're using it commercially you're probably deploying it on a server where you're not limited by the 24GB and you can just run llama 2 70b. The majority of people who want to run it locally on 24GB either want roleplay (so non commercial) or code (you have codellama)
- nabakin 3y agoLooks like they left out another model though. In the paper they mention a "Unnatural Code Llama" which wipes the floor with every other model/finetune on every benchmark except for slightly losing to Code Llama Python on MBPP pass@100 and slightly losing to GPT-4 on HumanEval pass@1 which is insane. Meta says later on that they aren't releasing it and give no explanation. I wonder why given how incredible it seems to be.
- up6w6 3y agoEven the 7B model of code llama seems to be competitive with Codex, the model behind copilot https://ai.meta.com/blog/code-llama-large-language-model-coding/ https://ai.meta.com/blog/code-llama-large-language-model-cod...
- ramesh31 3y ago>Even the 7B model of code llama seems to be competitive with Codex, the model behind copilot It's extremely good. I keep a terminal tab open with 7b running for all of my "how do I do this random thing" questions while coding. It's pretty much replaced Google/SO for me.
- coder543 3y agoYou've already downloaded and thoroughly tested the 7B parameter model of "code llama"? I'm skeptical.
- lddemi 3y agoLikely meta employee?
- MertsA 3y agoI've been using this or something similar internally for months and love it. The thing that gets downright spooky is the comments believe it or not. I'll have some method with a short variable name in a larger program and not only does it often suggest a pretty good snippet of code the comments will be correct and explain what the intent behind the code is. It's just a LLM but you really start to get the feeling the whole is greater than the sum of the parts.
- coder543 3y agoI just don’t understand how anyone is making practical use of local code completion models. Is there a VS Code extension that I’ve been unable to find? HuggingFace released one that is meant to use their service for inference, not your local GPU. The instruct version of code llama could certainly be run locally without trouble, and that’s interesting too, but I keep wanting to test out a local CoPilot alternative that uses these nice, new completion models.
- brucethemoose2 3y agoHere is the paper: https://ai.meta.com/research/publications/code-llama-open-foundation-models-for-code/ https://ai.meta.com/research/publications/code-llama-open-fo...
- ilaksh 3y agoBetween this, ideogram.ai (image generator which can spell, from former Google Imagen team member and others), and ChatGPT fine-tuning, this has been a truly epic week. I would argue that many teams will have to reevaluate their LLM strategy _again_ for the second time in a week.
- ShamelessC 3y agoDid ideogram release a checkpoint?
- ilaksh 3y agoI can't find any info or Discord or forum or anything. I think it's a closed service that they plan to sell to make money.
- astrange 3y agoSDXL and DeepFloyd can spell. It's more or less just a matter of having a good enough text encoder. I tried Ideogram yesterday and it felt too much like existing generators (base SD and Midjourney). DALLE2 actually has some interestingly different outputs, the problem is they never update it or fix the bad image quality.
- eurekin 3y agotheBloke cannot rest :)
- ynniv 3y agoEvery time a new model hits I'm waiting for his ggmls
- brucethemoose2 3y agoggml quantization is very easy with the official llama.cpp repo. Its quick and mostly dependency free, and you can pick the perfect size for your CPU/GPU pool. But don't get me wrong, TheBloke is a hero.
- ynniv 3y agoSome of the newer models have slightly different architectures, so he explains any differences and shows a llama.cpp invocation. Plus you can avoid pulling the larger dataset.
- brucethemoose2 3y agoYeah. Keeping up wkth the changes is madness, and those FP16 weights are huge.
- int_19h 3y agoWhile we're at it, the GGML file format has been deprecated in favor of GGUF. https://github.com/philpax/ggml/blob/gguf-spec/docs/gguf.md https://github.com/philpax/ggml/blob/gguf-spec/docs/gguf.md https://github.com/ggerganov/llama.cpp/pull/2398 https://github.com/ggerganov/llama.cpp/pull/2398
- regularfry 3y agoAs if by magic... https://huggingface.co/TheBloke/CodeLlama-13B-fp16 https://huggingface.co/TheBloke/CodeLlama-13B-fp16. Empty so still uploading right now, at a guess.
- MuffinFlavored 3y agoCan I feed this entire GitHub projects (of reasonable size) and get non-hallucinated up-to-date API refactoring recommendations?
- redox99 3y agoThe highlight IMO > The Code Llama models provide stable generations with up to 100,000 tokens of context. All models are trained on sequences of 16,000 tokens and show improvements on inputs with up to 100,000 tokens. Edit: Reading the paper, key retrieval accuracy really deteriorates after 16k tokens, so it remains to be seen how useful the 100k context is.
- brucethemoose2 3y agoDid Meta add scalable rope to the official implementation?
- snippyhollow 3y agoWe changed RoPE's theta from 10k to 1m and fine-tuned with 16k tokens long sequences.
- malwrar 3y agoCurious, what led you to adjusting the parameters this way? Also, have you guys experimented with ALiBi[1] which claims better extrapolative results than rotary positional encoding? [1]: https://arxiv.org/abs/2108.12409 https://arxiv.org/abs/2108.12409 (charts on page two if you’re skimming)
- ttul 3y agoUndoubtedly, they have tried ALiBi…
- deleted 3y ago[deleted]
- nabakin 3y agoLooks like they aren't releasing a pretty interesting model too. In the paper they mention a "Unnatural Code Llama" which wipes the floor with every other model/finetune on every benchmark except for slightly losing to Code Llama Python on MBPP pass@100 and slightly losing to GPT-4 on HumanEval pass@1 which is insane. Meta says later on that they aren't releasing it and give no explanation. I wonder why given how incredible it seems to be.
- ChatGTP 3y ago[flagged]
- ilaksh 3y agohttps://github.com/facebookresearch/codellama https://github.com/facebookresearch/codellama
- naillo 3y agoFeels like we're like a year away from local LLMs that can debug code reliably (via being hooked into console error output as well) which will be quite the exciting day.
- brucethemoose2 3y agoThat sounds like an interesting finetuning dataset. Imagine a database of "Here is the console error, here is the fix in the code" Maybe one could scrape git issues with console output and tagged commits.
- ilaksh 3y agoHave you tried Code Llama? How do you know it can't do it already? In my applications, GPT-4 connected to a VM or SQL engine can and does debug code when given error messages. "Reliably" is very subjective. The main problem I have seen is that it can be stubborn about trying to use outdated APIs and it's not easy to give it a search result with the correct API. But with a good web search and up to date APIs, it can do it. I'm interested to see general coding benchmarks for Code Llama versus GPT-4.
- tomr75 3y agoHave you tried giving up to date apis as context?
- jebarker 3y agoWhat does "GPT-4 connected to a VM or SQL engine" mean?
- ilaksh 3y agohttps://aidev.codes https://aidev.codes shows connected to VM.
- sumedh 3y ago> But with a good web search and up to date APIs, it can do it. How do you do that?
- 6stringmerc 3y agoSo it’s stubborn, stinks, bites and spits? No thanks, going back to Winamp.
- benvolio 3y ago>The Code Llama models provide stable generations with up to 100,000 tokens of context. Not a bad context window, but makes me wonder how embedded code models would pick that context when dealing with a codebase larger than 100K tokens. And this makes me further wonder if, when coding with such a tool (or at least a knowledge that they’re becoming more widely used and leaned on), are there some new considerations that we should be applying (or at least starting to think about) when programming? Perhaps having more or fewer comments, perhaps more terse and less readable code that would consume fewer tokens, perhaps different file structures, or even more deliberate naming conventions (like Hungarian notation but for code models) to facilitate searching or token pattern matching of some kind. Ultimately, in what ways could (or should) we adapt to make the most of these tools?
- brucethemoose2 3y agoThis sounds like a job for middleware. Condensing split code into a single huge file, shortening comments, removing whitespace and such can be done by a preprocessor for the llm.
- gabereiser 3y agoSo now we need an llmpack like we did webpack? Could it be smart enough to truncate comments, white space, etc?
- brucethemoose2 3y agoYou dont even need an llm for trimming whitespace, just a smart parser with language rules like ide code checkers already use. Existing llms are fine at summarizing comments, especially with language specific grammar constraints.
- gabereiser 3y agoMy point. We don’t need the middleware.
- 3y ago
- thehacker1234 3y ago[flagged]
- jasfi 3y agoNow we need code quality benchmarks comparing this against GPT-4 and other contenders.
- nick0garvey 3y agoThey show the benchmarks in the original post, a few pages down
- jasfi 3y agoThanks, I missed that somehow.
- braindead_in 3y agoThe 34b Python model is quite close to GPT4 on HumanEval pass@1. Small specialised models are catching up to GPT4 slowly. Why not train a 70b model though?
- binary132 3y agoI wonder whether org-ai-mode could easily support this.
- scriptsmith 3y agoHow are people using these local code models? I would much prefer using these in-context in an editor, but most of them seem to be deployed just in an instruction context. There's a lot of value to not having to context switch, or have a conversation. I see the GitHub copilot extensions gets a new release one every few days, so is it just that the way they're integrated is more complicated so not worth the effort?
- modeless 3y agohttp://cursor.sh http://cursor.sh integrates GPT-4 into vscode in a sensible way. Just swapping this in place of GPT-4 would likely work perfectly. Has anyone cloned the OpenAI HTTP API yet?
- fudged71 3y agoI was tasked with a massive project over the last month and I'm not sure I could have done it as fast as I have without Cursor. Also check out the Warp terminal replacement. Together it's a winning combo!
- lhl 3y agoLocalAI https://localai.io/ https://localai.io/ and LMStudio https://lmstudio.ai/ https://lmstudio.ai/ both have fairly complete OpenAI compatibility layers. llama-cpp-python has a FastAPI server as well: https://github.com/abetlen/llama-cpp-python/blob/main/llama_cpp/server/app.py https://github.com/abetlen/llama-cpp-python/blob/main/llama_... (as of this moment it hasn't merged GGUF update yet though)
- thewataccount 3y agoFor in-editor like copilot you can try this locally - https://github.com/smallcloudai/refact https://github.com/smallcloudai/refact This works well for me except the 15B+ don't run fast enough on a 4090 - hopefully exllama supports non-llama models, or maybe it'll support CodeLLaMa already I'm not sure. For general chat testing/usage this works pretty well with lots of options - https://github.com/oobabooga/text-generation-webui/ https://github.com/oobabooga/text-generation-webui/
- likenesstheft 3y agono more work soon?
- kypro 3y agoThe ability to work less historically has always came as a byproduct of individuals earning more per hour through productivity increases. The end goal of AI isn't to make your labour more productive, but to not need your labour at all. As your labour becomes less useful if anything you'll find you need to work more. At some point you may be as useful to the labour market as someone with 60 IQ today. At this point most of the world will become entirely financially dependent on the wealth redistribution of the few who own the AI companies producing all the wealth – assuming they take pity on you or there's something governments can actually do to force them to pay 90%+ tax rates, of course.
- likenesstheft 3y agoWhat?
- Dowwie 3y agoWhat did the fine tuning process consist of?
- the-alchemist 3y agoAnyone know if it supports Clojure?
- deleted 3y ago[deleted]
- 1024core 3y ago> Python, C++, Java, PHP, Typescript (Javascript), C#, and Bash What?!? No Befunge[0], Brainfuck or Perl?!? [0] https://en.wikipedia.org/wiki/Befunge https://en.wikipedia.org/wiki/Befunge /just kidding, of course!
- bick_nyers 3y agoAnyone know of a good plugin for the JetBrains IDE ecosystem (namely, PyCharm) that is CoPilot but with a local LLM?
- kateklink 3y agotry refact.ai they have plugin for JetBrains IDEs and support for local LLMs https://github.com/smallcloudai/refact/ https://github.com/smallcloudai/refact/
- rvnx 3y agoAmazing! It's great that Meta is making AI progress. In the meantime, we are still waiting for Google to show what they have (according to their research papers, they are beating others). > User: Write a loop in Python that displays the top 10 prime numbers. > Bard: Sorry I am just an AI, I can't help you with coding. > User: How to ask confirmation before deleting a file ? > Bard: To ask confirmation before deleting a file, just add -f to the rm command. (real cases)
- criley2 3y agoI don't get comments like this, we can all go and test Bard and see that what you're saying isn't true https://g.co/bard/share/95761dd6d45e https://g.co/bard/share/95761dd6d45e
- rvnx 3y agoWell look for yourself: https://g.co/bard/share/e8d14854ccab https://g.co/bard/share/e8d14854ccab The rm answer is now "hardcoded" (aka, manually entered by reviewers), the same with the prime or fibonnaci. This is why we both see the same code across different accounts (you can make the test if you are curious).
- criley2 3y agoOkay, so the entire point of the comment is "A current model which does well used to be bad!" With all due respect, is that a valuable thing to say? Isn't it true of them all?
- deleted 3y ago[deleted]
- rvnx 3y agoMhh not just about the past, you can see such in current answers from Bard. They are generally okayish, closer to "meh", than something outstanding. Yes the shell script solution is better, it doesn't give rm -f anymore, but is still somewhat closer to a bad solution instead of just giving rm -i. I'm just really happy and excited to see that a free-to-download and free-to-use model can beat a commercially-hosted offering. This is what has brought the most amazing projects (e.g. Stable Diffusion)
- dangerwill 3y agoIt's really sad how everyone here is fawning over tech that will destroy you own livelihoods. "AI won't take your job, those who use AI will" is purely short term, myopic thinking. These tools are not aimed to help workers, the end goal is to make it so you don't need to be an engineer to build software, just let the project manager or director describe the system they want and boom there it is. You can scream that this is progress all you want, and I'll grant you that these tools will greatly speed up the generation of code. But more code won't make any of these businesses provide better services to people, lower their prices, or pay workers more. They are just a means to keep money from flowing out of the hands of the C-Suite and investor classes. If software engineering becomes a solved problem then fine, we probably shouldn't continue to get paid huge salaries to write it anymore, but please stop acting like this is a better future for any of us normal folks.
- tomr75 3y agoImprove productivity, cheapen goods and services. Nature of technological advancement
- bbor 3y agoWe have three options, IMO: 1. As a species decide to never build another LLM, ever. 2. Change the path of society from the unequal, capitalist one it’s taken the last 2-300 years. 3. Give up I know which I believe in :). Do you disagree?
- pbhjpbhj 3y agoThe problem with 2 is that the people in power, and with immense wealth, remain there because of capitalism. They have the political power, and the resources, to enact the change ... but they also lose the most (unless you count altruism as gain, which of it were true the World would be so different). We could structure things so that LLM, and the generalised AIs to come, benefit the whole of society ... but we know that those with the power to make that happen want only to widen the poverty gap.
- 3y ago
- mdaniel 3y agoit looks like https://news.ycombinator.com/item?id=37248844 https://news.ycombinator.com/item?id=37248844 has gotten the traction at 295 points
- dang 3y agoMaybe we'll merge that one hither to split the karma.
- deleted 3y ago[deleted]
- andrewjl 3y agoWhat I found interesting in Meta's paper is the mention of HumanEval[1] and MBPP[2] as benchmarks for code quality. (Admittedly maybe they're well-known to those working in the field.) I haven't yet read the whole paper (nor have I looked at the benchmark docs which might very well cover this) but curious how these are designed to avoid issues with overfitting. My thinking here is that canned algorithm type problems common in software engineering interviews are probably over represented in the training data used for these models. Which might point to artificially better performance by LLMs versus their performance on more domain-specific type tasks they might be used for in day-to-day work. [1] https://github.com/openai/human-eval https://github.com/openai/human-eval [2] https://github.com/google-research/google-research/tree/master/mbpp https://github.com/google-research/google-research/tree/mast...
- lordnacho 3y agoCopilot has been working great for me thus far, but it's limited by its interface. It seems like it only knows how to make predictions for the next bit of text. Is anyone working on a code AI that can suggest refactorings? "You should pull these lines into a function, it's repetitive" "You should change this structure so it is easier to use" Etc
- artificialLimbs 3y agoI let mine generate whatever it likes, then add a comment below such as "# Refactor the above to foo.." Works fairly well at times.
- lordnacho 3y agoCan it suggest deletions? Just seems like I don't know how to use it.
- make3 3y agoThere's an instruct model in there, you can definitely use it for this, that's one of the objectives. An instruct model means that you can ask it to do what you want, including asking it to give you refactoring ideas from the code you will give it.
- regularfry 3y agoSounds like what's needed is a bit of tooling in the background consistently asking the LLM "How would you improve this code?" so you don't need to actually ask it.
- lordnacho 3y agoHow do I access it from my IDE? Jetbrains/VSCode?
- phillipcarter 3y agoSourceGraph Cody is going in that direction, as is Copilot Chat. But it's still early days. I don't think there's anything robust here yet.
- natch 3y agoWhy wouldn’t they provide a hosted version? Seems like a no brainer… they have the money, the hardware, the bandwidth, the people to build support for it, and they could design the experience and gather more learning data about usage in the initial stages, while putting a dent in ChatGPT commercial prospects, and all while still letting others host and use it elsewhere. I don’t get it. Maybe it was just the fastest option?
- redox99 3y agoProbably the researchers at meta are only interested in research, and productionizing this would be up to other teams.
- natch 3y agoBut Yann LeCun seems to think the safety problems of eventual AGI will be solved somehow. Nobody is saying this model is AGI obviously. But this would be an entry point into researching one small sliver of the alignment problem. If you follow my thinking, it’s odd that he professes confidence that AI safety is a non issue, yet from this he seems to want no part in understanding it. I realize their research interest may just be the optimization / mathy research… that’s their prerogative but it’s odd imho.
- ShamelessC 3y agoIt’s not that odd and I think you’re overestimating the importance of user submitted data for the purposes of alignment research. In particular because it’s more liability for them to try to be responsible for outputs. Really though, this way they get a bunch of free work from volunteers in open source/ML communities.
- natch 3y agoYes sounds like a reasonable explanation, thanks.
- marcopicentini 3y agoIt's just a matter of time that Microsoft will integrate it into VSCode.
- Palmik 3y agoThe best model, Unnatural Code Llama, is not released. Likely because it's trained on GPT4 based data, and might violate OpenAI TOS, because as per the "Unnatural" paper [1], the "unnatural" data is generated with the help of some LLM -- and you would want to use as good of an LLM as possible. [1] https://arxiv.org/pdf/2212.09689.pdf https://arxiv.org/pdf/2212.09689.pdf
- redox99 3y agoThe good thing is that if it's only finetuned on 15k instructions, we should see a community made model like that very soon.
- Draiken 3y agoAs a complete noob at actually running these models, what kind of hardware are we talking here? Couldn't pick that up from the README. I absolutely love the idea of using one of these models without having to upload my source code to a tech giant.
- liuliu 3y ago34B should be able to run on 24GiB consumer graphics card, or 32GiB Mac (M1 / M2 chips) with quantization (5~6bit) (and 7B should be able to run on your smart toaster).
- epolanski 3y agoAre there cloud offerings to run those models on somebody's else computer? Any "eli5" tutorial on how to do so, if so? I want to give these models a run but I have no powerful GPU to run them on so don't know where to start.
- redox99 3y agoOn runpod there is a TheBloke template with everything set up for you. An A6000 is good enough to run 70b 4bit.
- kordlessagain 3y agoI started something here about this: https://news.ycombinator.com/item?id=37121384 https://news.ycombinator.com/item?id=37121384
- RossBencina 3y agorunpod, togethercomputer, replicate. Matthew Berman has a tutorial on YT showing how to use TheBloke's docker containers on runpod. Sam Witteveen has done videos on together and replicate, they both offer cloud-hosted LLM inference as a service.
- dangelov 3y agoI've used Ollama to run Llama 2 (all variants) on my 2020 Intel MacBook Pro - it's incredibly easy. You just install the app and run a couple of shell commands. I'm guessing soon-ish this model will be available too and then you'd be able to use it with the Continue VS Code extension. Edited to add: Though somewhat slow, swap seems to have been a good enough replacement for not having the loads of RAM required. Ollama says "32 GB to run the 13B models", but I'm running the llama2:13b model on a 16 GB MBP.
- bryanlyon 3y agoLlama is a very cool language model, it being used for coding was all but inevitable. I especially love it being released open for everyone. I do wonder about how much use it'll get, seeing as running a heavy language model on local hardware is kinda unlikely for most developers. Not everyone is runnning a system powerful enough to equip big AIs like this. I also doubt that companies are going to set up large AIs for their devs. It's just a weird positioning.
- outside1234 3y ago... "seeing as running a heavy language model on local hardware is kinda unlikely for most developers" for now it is :) but with quantization advances etc. it is not hard to see the trajectory.
- ctoth 3y agoAs we all know, computers stay the same and rarely improve.
- int_19h 3y ago12Gb of VRAM lets you run 13B models (4-bit quantized) with reasonable speed, and can be had for under $300 if you go for previous-generation NVidia hardware. Plenty of developers around with M1 and M2 Macs, as well.
- syntaxing 3y agoTheBloke doesn’t joke around [1]. I’m guessing we’ll have the quantized ones by the end of the day. I’m super excited to use the 34B Python 4 bit quantized one that should just fit on a 3090. [1] https://huggingface.co/TheBloke/CodeLlama-13B-Python-fp16 https://huggingface.co/TheBloke/CodeLlama-13B-Python-fp16
- suyash 3y agocan it be quantised further so it can run locally on a normal laptop of a developer?
- syntaxing 3y ago“Normal laptop” is kind of hard to gauge but if you have a M series MacBook with 16GB+ RAM, you will be able to run 7B comfortably and 13B but stretching your RAM (cause of the unified RAM) at 4 bit quantization. These go all the way down to 2 bit but I personally I find the model noticeably deteriorate anything below 4 bit. You can see how much (V)RAM you need here [1]. [1] https://github.com/ggerganov/llama.cpp#quantization https://github.com/ggerganov/llama.cpp#quantization
- totallywrong 3y agoJust started playing with this, there's a tool called ollama that runs Llama2 13B on my 16GB M1 Pro really smoothly with zero config.
- stuckinhell 3y agoWhat kind of cpu/gpu power do you need for quantization or these new gguf formats ?
- syntaxing 3y agoI haven’t quantized these myself since TheBloke has been the main provider for all the quantized models. But when I did a 8 bit quantization to see how it compares to the transformers library load_in_8bit 4 months ago(?), it didn’t use my GPU but loaded each shard into the RAM during the conversion. I had an old 4C/8T CPU and the conversion took like 30 mins for a 13B.
- mercurialsolo 3y agoIs there a version of this on replicate yet?
- gdcbe 3y agoIs there somewhere docs to show you how to run this on your local machine and can you make it port it a script between languages? Gpt4 can do that pretty well but its context is too small for advanced purposes.
- born-jre 3y ago34B is grouped query attention, right? Does that make it the smallest model with grouped attention? I can see some people fine-tuning it again for general propose instruct.
- e12e 3y agoCurious if there are projects to enable working with these things self-hosted, tuned to a git repo as context on the cli, like a Unix filter - or with editors like vim? (I'd love to use this with Helix) I see both vscode and netbeans have a concept of "inference URL" - are there any efforts like language server (lsp) - but for inference?
- jmorgan 3y agoTo run Code Llama locally, the 7B parameter quantized version can be downloaded and run with the open-source tool Ollama: https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama ollama run codellama "write a python function to add two numbers" More models coming soon (completion, python and more parameter counts)
- WaitWaitWha 3y agoCan someone point me to a ELI5 sequence of steps that shows how someone can install and use LLMs locally and in some way, functionally? Asking for purposes of educating non-technologists.
- Patrick_Devine 3y agoThere are several different ways, but the easiest way in my (clearly biased) opinion is just got to ollama.ai, download it, and start playing around. It works out of the box w/ newer Macs, but there are versions for Linux and Windows in the works.
- deleted 3y ago[deleted]
- daemonologist 3y agoWorks nearly out of the box with llama.cpp, which makes it easy to try locally: https://github.com/ggerganov/llama.cpp/issues/2766 https://github.com/ggerganov/llama.cpp/issues/2766 Here's some output from q4_0 quantization of CodeLlama-7b-Python (first four lines are the prompt): # prints the first ten prime numbers def print_primes(): i = 2 num_printed = 0 # end of prompt while num_printed < 10: if is_prime(i): print(i) num_printed += 1 i += 1 def is_prime(n): i = 2 while i * i <= n: if n % i == 0: return False i += 1 return True def main(): print_primes() if __name__ == '__main__': main() It will be interesting to see how the larger models perform, especially after community tuning and with better context/prompting.
- blibble 3y agoI'd fail an interview candidate that suggested adding 1 each time for subsequent prime testing
- droopyEyelids 3y agoThe simple-to-understand, greedy algorithm is always the correct first choice till you have to deal with a constraint.
- blibble 3y agoit's not that though, there's several other typical optimisations in there just not the super obvious one that demonstrates extremely basic understanding of what a prime number is
- jpeterson 3y agoHaving "extremely basic understanding" of prime numbers immediately at one's command is important for approximately 0% of software engineering jobs. If you instant-fail a candidate for this, it says a lot more about you and your organization than the candidate.
- praveenhm 3y agowhich is the best model for coding right now, GPT4/copilot/phind ?
- bracketslash 3y agoSo uhh…how does one go about using it?
- gorbypark 3y agoI can't wait for some models fine tuned on other languages. I'm not a Python developer, so I downloaded the 13B-instruct variant (4 bit quantized Q4_K_M) and it's pretty bad at doing javascript. I asked it to write me a basic React Native component that has a name prop and displays that name. Once it returned a regular React component, and when I asked it to make sure it uses React Native components, it said sure and outputted a bunch of random CSS and an HTML file that was initializing a React project. It might be the quantization or my lacklustre prompting skills affecting it, though. To be fair I did get it to output a little bit of useful code after trying a few times.
- rafaelero 3y agoThose charts remind me just how insanely good GPT-4 is. It's almost 5 months since its release and I am still at awe with its capabilities. The way it helps with coding is just crazy.
- 1024core 3y agoIf GPT-4's accuracy is 67% and this is 54%, how can these guys claim to be SOTA?
- Someone1234 3y agoBusiness opportunity: I'd pay money for NICE desktop software that can run all these different models (non-subscription, "2-year updates included, then discount pricing" modal perhaps). My wishlist: - Easy plug & play model installation, and trivial to change which model once installed. - Runs a local web server, so I can interact with it via any browser - Ability to feed a model a document or multiple documents and be able to ask questions about them (or build a database of some kind?). - Absolute privacy guarantees. Nothing goes off-machine from my prompt/responses (USP over existing cloud/online ones). Routine license/update checks are fine though. I'm not trying to throw shade at the existing ways to running LLMs locally, just saying there may be room for an OPTIONAL commercial piece of software in this space. Most of them are designed for academics to do academic things. I am talking about a turn-key piece of software for everyone else that can give you an "almost" ChatGPT or "almost" CoPilot-like experience for a one time fee that you can feed sensitive private information to.
- jmorgan 3y agoA few folks and I have been working on an open-source tool that does some of this (and hopefully more soon!) https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama There's a "PrivateGPT" example in there that is similar to your third point above: https://github.com/jmorganca/ollama/tree/main/examples/privategpt https://github.com/jmorganca/ollama/tree/main/examples/priva... Would love to know your thoughts
- SubiculumCode 3y ago[flagged]
- luma 3y agoI'd love to test this out as soon as you get Linux or Windows support going!
- appel 3y agoMe too! I starred the repo and am watching releases, excited to try it.
- jtwaleson 3y agoThis is probably a stupid question, but would it be possible to use these models to rate existing code and point to possible problems, rather than generating new code? That would be extremely useful to some use cases I'm working on.
- WhitneyLand 3y agoHow much am I’m missing out on with tools like this or code pilot, compared to using GPT-4? I guess since Xcode doesn’t have a good plug-in architecture for this I began experimenting more with a chat interface. So far gpt-4 has seemed quite useful for generating code, reviewing code for certain problems, etc.
- citruscomputing 3y agoEditor plugins are fantastic about completing based on a pattern. That's the main thing you're missing out on imo - it's worth it to hit tab, but not to copy/paste and say "finish this line for me, it looks almost like the one above." There's also the real-time aspect where you can see that it's wrong via the virtual text, type a few characters, then it gets what you're doing and you can tab complete the rest of the line. It's faster to converse with when you don't have to actually have a conversation, if that makes sense? The feedback loop is much shorter and doesn't require natural language, or nearly as much context switching.
- KaiserPro 3y agoThis is great for asking questions like "how do I do x with y" and this code <<some code>> isn't working, whats wrong? Much faster that googling, or a great basis for forming a more accurate google search. Where its a bit shit is when its used to provide auto suggest. It hallucinates plausible sounding functions/names, which for me personally are hard to stop if they are wrong (I suspect that's a function of the plugin)
- SubiculumCode 3y agohallucinations can be resuces by incorporating 'retrieval automated generation' , RAG, on the front end. likely function library defs could be automagically entered as prompt/memory inputs.
- pmarreck 3y agoI want "safety" to be opt-in due to the inaccuracy it introduces. I don't want to pay that tax just because someone is afraid I can ask it how to make a bomb when I can just Google that and get pretty close to the same answer already, and I certainly don't care about being offended by its answers.
- dontupvoteme 3y agoDid people *really* think only artists would be losing their jobs to AI?
- jrh3 3y agolol... Python for Dummies (TM)
- akulbe 3y agoRandom tangential question given this is about llama, but how do you get llama.cpp or kobold (or whatever tool you use) to make use of multiple GPUs if you don't have NVlink in place? I got a bridge, but it was the wrong size. Thanks, in advance.
- ai_g0rl 3y agothis is cool, https://labs.perplexity.ai/ https://labs.perplexity.ai/ has been my favorite way to play w these models so far
- nothrowaways 3y agoKudos to the team at FB.
- TheRealClay 3y agoAnyone know of a docker image that provides an HTTP API interface to Llama? I'm looking for a super simple sort of 'drop-in' solution like that which I can add to my web stack, to enable LLM in my web app.
- nodja 3y agohttps://github.com/abetlen/llama-cpp-python https://github.com/abetlen/llama-cpp-python has a web server mode that replicates openai's API iirc and the readme shows it has docker builds already.
- TheRealClay 3y agoThanks! As someone just getting started, I really appreciate the tip!
- dchuk 3y agoGiven this can produce code when prompted, could it also be used to interpret html from a crawler and then be used to scrape arbitrary URLs and extract structured attributes? Basically like MarkupLM but with massively more token context?
- stevofolife 3y agoAlso curious about this. There must be a better way to scrape using LLM.
- awwaiid 3y agoI want to see (more) code models trained on git diffs
- jerrygoyal 3y agowhat is the cutoff knowledge of it? Also, what is the cheapest way to use it if I'm building a commercial tool on top of it?
- m00nsome 3y agoWhy do they not release the unnatural Variant of the model? According to the paper it beats all of the other variants and seems to be close to GPT-4.
- KingOfCoders 3y agoAny performance tests? (e.G. tokens/s on a 4090?)
- robertnishihara 3y agoIf you want to try out Code Llama, you can query it on Anyscale Endpoints (this is an LLM inference API we're working on here at Anyscale). https://app.endpoints.anyscale.com/ https://app.endpoints.anyscale.com/
- pelorat 3y agoTo bad most models focus on Python, it's not a popular language here in Europe (for anything).
- Havoc 3y agoWhat’s Europe using for machine learning?
- RobKohr 3y agoNow it just needs a vscode plugin to replace copilot.