6 ms·
Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
- parmesean 2y ago13.8 epochs of the benchmarks?
- xeckr 2y agoLooks like it punches way above its weight(s). How far are we from running a GPT-3/GPT-4 level LLM on regular consumer hardware, like a MacBook Pro?
- bloomingkales 2y agoM4 Mac mini 16gb for $500. It's literally an inferencing block (small too, fits in my palm). I feel like the whole world needs one.
- alganet 2y ago> inferencing block Did you mean _external gpu_? Choose any 12GB or more video card with GDDR6 or superior and you'll have at least double the performance of a base m4 mini. The base model is almost an older generation. Thunderbolt 4 instead of 5, slower bandwidths, slower SSDs.
- kgwgk 2y ago> you'll have at least double the performance of a base m4 mini For $500 all included?
- alganet 2y agoThe base mini is 599. Here's a config for around the same price. All brand new parts for 573. You can spend the difference improving any part you wish, or maybe get an used 3060 and go AM5 instead (Ryzen 8400F). Both paths are upgradeable. https://pcpartpicker.com/list/ftK8rM https://pcpartpicker.com/list/ftK8rM Double the LLM performance. Half the desktop performance. But you can use both at the same time. Your computer will not slow down when running inference.
- bloomingkales 2y agoThat’s a really nice build.
- alganet 2y agoAnother possible build is to use a mini-pc and M.2 connections You'll need a mini-pc with two M.2 slots, like this: https://www.amazon.com/Beelink-SER7-7840HS-Computer-Display/dp/B0CQT9N951 https://www.amazon.com/Beelink-SER7-7840HS-Computer-Display/... And a riser like this: https://www.amazon.com/CERRXIAN-Graphics-Left-PCI-Express-Extension/dp/B0D9651WSN https://www.amazon.com/CERRXIAN-Graphics-Left-PCI-Express-Ex... And some courage to open it and rig the stuff in. Then you can plug a GPU on it. It should have decent load times. Better than an eGPU, worse than the AM4 desktop build, fast enough to beat the M4 (once the data is in the GPU, it doesn't matter). It makes for a very portable setup. I haven't built it, but I think it's a reasonable LLM choice comparable to the M4 in speed and portability while still being upgradable. Edit: and you'll need an external power supply of at least 400W:)
- lappa 2y agoIt's easy to argue that Llama-3.3 8B performs better than GPT-3.5. Compare their benchmarks, and try the two side-by-side. Phi-4 is yet another step towards a small, open, GPT-4 level model. I think we're getting quite close. Check the benchmarks comparing to GPT-4o on the first page of their technical report if you haven't already https://arxiv.org/pdf/2412.08905 https://arxiv.org/pdf/2412.08905
- vulcanash999 2y agoDid you mean Llama-3.1 8B? Llama 3.3 currently only has a 70B model as far as I’m aware.
- anon373839 2y agoWe’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I have an M3 Max and I can’t say that models of this size run at comfortable speeds. They’re a bit sluggish.
- noman-land 2y agoThe coolness of local LLMs is THE only reason I am sadly eyeing upgrading from M1 64GB to M4/5 128+GB.
- Terretta 2y agoCompare performance on various Macs here as it gets updated: https://github.com/ggerganov/llama.cpp/discussions/4167 https://github.com/ggerganov/llama.cpp/discussions/4167 OMM, Llama 3.3 70B runs at ~7 text generation tokens per second on Macbook Pro Max 128GB, while generating GPT-4 feeling text with more in depth responses and fewer bullets. Llama 3.3 70B also doesn't fight the system prompt, it leans in. Consider e.g. LM Studio (0.3.5 or newer) for a Metal (MLX) centered UI, include MLX in your search term when downloading models. Also, do not scrimp on the storage. At 60GB - 100GB per model, it takes a day of experimentation to use 2.5TB of storage in your model cache. And remember to exclude that path from your TimeMachine backups.
- noman-land 2y agoThank you for all the tips! I'd probably go 128GB 8TB because of masochism. Curious, what makes so many of the M4s in the red currently.
- vessenes 2y agoIt's all memory bandwidth related -- what's slow is loading these models into memory, basically. The last die from Apple with all the channels was the M2 Ultra, and I bet that's what tops those leader boards. M4 has not had a Max or an Ultra release yet; when it does (and it seems likely it will), those will be the ones to get.
- simonw 2y agoWe're there. Llama 3.3 70B is GPT-4 level and runs on my 64GB MacBook Pro: https://simonwillison.net/2024/Dec/9/llama-33-70b/ https://simonwillison.net/2024/Dec/9/llama-33-70b/ The Qwen2 models that run on my MacBook Pro are GPT-4 level too.
- BoorishBears 2y agoSaying these models are at GPT-4 level is setting anyone who doesn't place special value on the local aspect up for disappointment. Some people do place value on running locally, and I'm not against then for it, but realistically no 70B class model has the amount of general knowledge or understanding of nuance as any recent GPT-4 checkpoint. That being said these models are still very strong compared to what we had a year ago and capable of useful work
- simonw 2y agoI said GPT-4, not GPT-4o. I'm talking about a model that feels equivalent to the GPT-4 we were using in March of 2023.
- int_19h 2y agoI remember using GPT-4 when it first dropped to get a feeling of its capabilities, and no, I wouldn't say that llama-3.3-70b is comparable. At the end of the day, there's only so much you can cram into any given number of parameters, regardless of what any artificial benchmark says.
- simonw 2y agoI envy your memory.
- BoorishBears 2y agoYou're free to intentionally miss their point, does them no good.
- n144q 2y ago
- refulgentis 2y agoWe're there, Llama 3.1 8B beats Gemini Advanced for $20/month. Telosnex with llama 3.1 8b GGUF from bartowski. https://telosnex.com/compare/ https://telosnex.com/compare/ (How!? tl;dr: I assume Google is sandbagging and hasn't updated the underlying Gemini)
- deleted 2y ago[deleted]
- ActorNightly 2y agoWhy would you want to though? You already can get free access to large LLMs and nobody is doing anything groundbreaking with them.
- jckahn 2y agoI only use local, open source LLMs because I don’t trust cloud-based LLM hosts with my data. I also don’t want to build a dependence on proprietary technology.
- simonw 2y agoThe most interesting thing about this is the way it was trained using synthetic data, which is described in quite a bit of detail in the technical report: https://arxiv.org/abs/2412.08905 https://arxiv.org/abs/2412.08905 Microsoft haven't officially released the weights yet but there are unofficial GGUFs up on Hugging Face already. I tried this one: https://huggingface.co/matteogeniaccio/phi-4/tree/main https://huggingface.co/matteogeniaccio/phi-4/tree/main I got it working with my LLM tool like this: llm install llm-gguf llm gguf download-model https://huggingface.co/matteogeniaccio/phi-4/resolve/main/phi-4-Q4_K_M.gguf llm chat -m gguf/phi-4-Q4_K_M Here are some initial transcripts: https://gist.github.com/simonw/0235fd9f8c7809d0ae078495dd630b67 https://gist.github.com/simonw/0235fd9f8c7809d0ae078495dd630... More of my notes on Phi-4 here: https://simonwillison.net/2024/Dec/15/phi-4-technical-report/ https://simonwillison.net/2024/Dec/15/phi-4-technical-report...
- syntaxing 2y agoWow, those responses are better than I expected. Part of me was expecting terrible responses since Phi-3 was amazing on paper too but terrible in practice.
- refulgentis 2y agoOne of the funniest tech subplots in recent memory. TL;DR it was nigh-impossible to get it to emit the proper "end of message" token. (IMHO the chat training was too rushed). So all the local LLM apps tried silently hacking around it. The funny thing to me was no one would say it out loud. Field isn't very consumer friendly, yet.
- TeMPOraL 2y agoSpeaking of, I wonder if and how many of the existing frontends, interfaces and support packages that generalize over multiple LLMs, and include Anthropic, actually know how to prompt it correctly. Seems like most developers missed the memo on https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/use-xml-tags https://docs.anthropic.com/en/docs/build-with-claude/prompt-..., and I regularly end up in situation in which I wish they gave more minute control on how the request is assembled (proprietary), and/or am considering gutting the app/library myself (OSS; looking at you, Aider), just to have file uploads, or tools, or whatever other smarts the app/library does, encoded in a way that uses Claude to its full potential. I sometimes wonder how many other model or vendor-specific improvements there are, that are missed by third-party tools despite being well-documented by the vendors.
- thot_experiment 2y agoFor prompt adherence it still fails on tasks that Gemma2 27b nails every time. I haven't been impressed with any of the Phi family of models. The large context is very nice, though Gemma2 plays very well with self-extend.
- CuriousCosmic 2y agoYeah they mention this in the weaknesses section. > While phi-4 demonstrates relatively strong performance in answering questions and performing reasoning tasks, it is less proficient at rigorously following detailed instructions, particularly those involving specific formatting requirements.
- thot_experiment 2y agoAh good catch, I am forever cursed in my preference for snake over camel.
- impossiblefork 2y agoIt's a much smaller model though. I think the point is more the demonstration that such a small model can have such good performance than any actual usefulness.
- magicalhippo 2y agoGemma2 9B has significantly better prompt adherence than Llama 3.1 8B in my experience. I've just assumed it's down to how it was trained, but no expert.
- travisgriggs 2y agoWhere have I been? What is a “small” language model? Wikipedia just talks about LLMs. Is this a sort of spectrum? Are there medium language models? Or is it a more nuanced classifier?
- dboreham 2y agoThere are all sizes of models from a few GB to hundreds of GB. Small presumably means small enough to run on end-user hardware.
- narag 2y ago7B vs 70B parameters... I think. The small ones fit in the memory of consumer grade cards. That's what I more or less know (waiting for my new computer to arrive this week)
- agnishom 2y agoHow many parameters did ChatGPT have in Dec 2022 when it first broke into mainstream news?
- simonw 2y agoI don't think that's ever been shared, but it's predecessor GPT-3 Da Vinci was 175B. One of the most exciting trends of the past year has been models getting dramatically smaller while maintaining similar levels of capability.
- reissbaker 2y agoGPT-3 had 175B, and the original ChatGPT was probably just a GPT-3 finetune (although they called it gpt-3.5, so it could have been different). However, it was severely undertrained. Llama-3.1-8B is better in most ways than the original ChatGPT; a well-trained ~70B usually feels GPT-4-level. The latest Llama release, llama-3.3-70b, goes toe-to-toe even with much larger models (albeit is bad at coding, like all Llama models so far; it's not inherent to the size, since Qwen is good, so I'm hoping the Llama 4 series is trained on more coding tokens).
- _ea1k 2y agoI really like the ~3B param version of phi-3. It wasn't very powerful and overused memory, but was surprisingly strong for such a small model. I'm not sure how I can be impressed by a 14B Phi-4. That isn't really small any more, and I doubt it will be significantly better than llama 3 or Mistral at this point. Maybe that will be wrong, but I don't have high hopes.
- excerionsforte 2y agoLooks like someone converted it for Ollama use already: https://ollama.com/vanilj/Phi-4 https://ollama.com/vanilj/Phi-4
- accrual 2y agoI've had great success with quantized Phi-4 12B and Ollama so far. It's as fast as Llama 3.1 8B but the results have been (subjectively) higher quality. I copy/pasted some past requests into Phi-4 and found the answers were generally better.
- ai_biden 2y agoI'm not too excited by Phi-4 benchmark results - It is#BenchmarkInflation. Microsoft Research just dropped Phi-4 14B, an open-source model that’s turning heads. It claims to rival Llama 3.3 70B with a fraction of the parameters — 5x fewer, to be exact. What’s the secret? Synthetic data. -> Higher quality, Less misinformation, More diversity But the Phi models always have great benchmark scores, but they always disappoint me in real-world use cases. Phi series is famous for to be trained on benchmarks. I tried again with the hashtag#phi4 through Ollama - but its not satisfactory. To me, at the moment - IFEval is the most important llm benchmark. But look the smart business strategy of Microsoft: have unlimited access to gpt-4 the input prompt it to generate 30B tokens train a 1B parameter model call it phi-1 show benchmarks beating models 10x the size never release the data never detail how to generate the data( this time they told in very high level) claim victory over small models
- mupuff1234 2y agoSo we moved from "reasoning" to "complex reasoning". I wonder what will be next month's buzzphrase.
- TeMPOraL 2y ago> So we moved from "reasoning" to "complex reasoning". Only from the perspective of those still complaining about the use of the term "reasoning", who now find themselves left behind as the world has moved on. For everyone else, the phrasing change perfectly fits the technological change.
- HarHarVeryFunny 2y agoReasoning basically means multi-step prediction, but to be general the reasoner also needs to be able to: 1) Realize when it's reached an impasse, then backtrack and explore alternatives 2) Recognize when no further progress towards the goal appears possible, and switch from exploiting existing knowledge to exploring/acquiring new knowledge to attempt to proceed. An LLM has limited agency, but could for example ask a question or do a web search. In either case, prediction failure needs to be treated as a learning signal so the same mistake isn't repeated, and when new knowledge is acquired that needs to be remembered. In both cases this learning would need to persist beyond the current context in order to be something that the LLM can build on in the future - e.g. to acquire a job skill that may take a lot of experience/experimentation to master. It doesn't matter what you call it (basic or advanced), but it seems that current attempts at adding reasoning to LLMs (e.g. GPT-o1) are based around 1), a search-like strategy, and learning is in-context and ephemeral. General animal-like reasoning needs to also support 2) - resolving impasses by targeted new knowledge acquisition (and/or just curiosity-driven experimentation), as well as continual learning.
- criddell 2y agoIf you graded humanity on their reasoning ability, I wonder where these models would score? I think once they get to about the 85th percentile, we could upgrade the phrase to advanced reasoning. I'm roughly equating it with the percentage of the US population with at least a master's degree.
- zurfer 2y agoModel releases without comprehensive coverage of benchmarks make me deeply skeptical. The worst was the gpt4o update in November. Basically a 2 liner on what it is better at and in reality it regressed in multiple benchmarks. Here we just get MMLU, which is widely known to be saturated and knowing they trained on synthetic data, we have no idea how much "weight" was given to having MMLU like training data. Benchmarks are not perfect, but they give me context to build upon. --- edit: the benchmarks are covered in the paper: https://arxiv.org/pdf/2412.08905 https://arxiv.org/pdf/2412.08905
- PoignardAzur 2y agoSaying that a 14B model is "small" feels a little silly at this point. I guess it doesn't require a high-end graphics card?
- liminal 2y agoIs 14B parameters still considered small?