31 ms·
MCP in LM Studio
- chisleu 1y agoJust ordered a $12k mac studio w/ 512GB of integrated RAM. Can't wait for it to arrive and crank up LM Studio. It's literally the first install. I'm going to download it with safari. LM Studio is newish, and it's not a perfect interface yet, but it's fantastic at what it does which is bring local LLMs to the masses w/o them having to know much. There is another project that people should be aware of: https://github.com/exo-explore/exo https://github.com/exo-explore/exo Exo is this radically cool tool that automatically clusters all hosts on your network running Exo and uses their combined GPUs for increased throughput. Like HPC environments, you are going to need ultra fast interconnects, but it's just IP based.
- dchest 1y agoI'm using it on MacBook Air M1 / 8 GB RAM with Qwen3-4B to generate summaries and tags for my vibe-coded Bloomberg Terminal-style RSS reader :-) It works fine (the laptop gets hot and slow, but fine). Probably should just use llama.cpp server/ollama and not waste a gig of memory on Electron, but I like GUIs.
- minimaxir 1y ago8 GB of RAM with local LLMs in general is iffy: a 8-bit quantized Qwen3-4B is 4.2GB on disk and likely more in memory. 16 GB is usually the minimum to be able to run decent models without compromising on heavy quantization.
- hnuser123456 1y agoBut 8GB of Apple RAM is 16GB of normal RAM. https://www.pcgamer.com/apple-vp-says-8gb-ram-on-a-macbook-pro-analogous-to-16gb-ram-on-a-pc-we-just-happen-to-be-able-to-use-it-much-more-efficiently/ https://www.pcgamer.com/apple-vp-says-8gb-ram-on-a-macbook-p...
- arrty88 1y agoI concur. I just upgraded from m1 air with 8gb to m4 with 24gb. Excited to run bigger models.
- diggan 1y ago> m4 with 24gb Wow, that is probably analogous to 48GB on other systems then, if we were to ask an Apple VP?
- vntok 1y agoNot sure what Apple VPs have to do with the tech but yeah, pretty much any core engineer you ask at Apple will tell you this. Here is a nice article with some info about what memory compression is and how it works: https://arstechnica.com/gadgets/2013/10/os-x-10-9/#page-17 https://arstechnica.com/gadgets/2013/10/os-x-10-9/#page-17 It's been a hard technical problem but is pretty much solved by now since its first debut in 2012-2013.
- pxc 1y agoI've heard good things about how macOS handles memory relative to other operating systems. But Linux and Windows both have memory compression nowadays. So the claim is then not that memory compression makes your RAM twice as effective, but that macOS' memory compression is twice as good as the real and existing memory compression available on other operating systems. Doesn't such a claim... need stronger evidence?
- minimaxir 1y agoInterestingly it was AI (Apple Intelligence) that was the primary reason Apple abandoned that hedge.
- dchest 1y agoIt's 4-bit quantized (Q4_K_M, 2.5 GB) and still works well for this task. It's amazing. I've been running various small models on this 8 GB Air since the first Llama and GPT-J, and they improved so much! macOS virtual memory works well on swapping in and out stuff to SSD.
- karmakaze 1y agoNice. Ironically well suited for non-Apple Intelligence.
- incognito124 1y ago> I'm going to download it with Safari Oof you were NOT joking
- noman-land 1y agoSafari to download LM Studio. LM Studio to download models. Models to download Firefox.
- teaearlgraycold 1y agoThe modern ninite
- sneak 1y agoI already got one of these. I’m spoiled by Claude 4 Opus; local LLMs are slower and lower quality. I haven’t been using it much. All it has on it is LM Studio, Ollama, and Stats.app. > Can't wait for it to arrive and crank up LM Studio. It's literally the first install. I'm going to download it with safari. lol, yup. same.
- chisleu 1y agoYup, I'm spoiled by Claude 3.7 Sonnet right now. I had to stop using opus for plan mode in my Agent because it is just so expensive. I'm using Gemini 2.5 pro for that now. I'm considering ordering one of these today: https://www.newegg.com/p/N82E16816139451?Item=N82E16816139451&SoldByNewegg=1 https://www.newegg.com/p/N82E16816139451?Item=N82E1681613945... It looks like it will hold 5 GPUs with a single slot open for infiniband Then local models might be lower quality, but it won't be slow! :)
- kristopolous 1y agoThe GPUs are the hard things to find unless you want to pay like 50% markup
- sneak 1y agoThat’s just what they cost; MSRP is irrelevant. They’re not hard to find, they’re just expensive.
- evo_9 1y agoI was using Claude 3.7 exclusively for coding, but it sure seems like it got worse suddenly about 2–3 weeks back. It went from writing pretty solid code I had to make only minor changes to, to being completely off its rails, altering files unrelated to my prompt, undoing fixes from the same conversation, reinventing db access and ignoring existing coding 'standards' established in the existing codebase. Became so untrustworthy I finally gave OpenAi O3 a try and honestly, I was pretty surprised how solid it has been. I've been using o3 since, and I find it generally does exactly what I ask, esp if you have a well established project with plenty of code for it to reference. Just wondering if Claude 3.7 has seemed differently lately for anyone else? Was my go to for several months, and I'm no fan of OpenAI, but o3 has been rock solid.
- teaearlgraycold 1y agoWhat are you going to do with the LLMs you run?
- chisleu 1y agoCurrently I'm using gemini 2.5 and claude 3.7 sonnet for coding tasks. I'm interested in using models for code generation, but I'm not expecting much in that regard. I'm planning to attempt fine tuning open source models on certain tool sets, especially MCP tools.
- prettyblocks 1y agoI've been using openwebui and am pretty happy with it. Why do you like lm studio more?
- truemotive 1y agoOpen WebUI can leverage the built in web server in LM Studio, just FYI in case you thought it was primarily a chat interface.
- prophesi 1y agoNot OP, but with LM Studio I get a chat interface out-of-the-box for local models, while with openwebui I'd need to configure it to point to an OpenAI API-compatible server (like LM Studio). It can also help determine which models will work well with your hardware. LM Studio isn't FOSS though. I did enjoy hooking up OpenWebUI to Firefox's experimental AI Chatbot. (browser.ml.chat.hideLocalhost to false, browser.ml.chat.provider to localhost:${openwebui-port})
- s1mplicissimus 1y agoi recently tried openwebui but it was so painful to get it to run with local model. that "first run experience" of lm studio is pretty fire in comparison. can't really talk about actually working with it though, still waiting for the 8GB download
- prettyblocks 1y agoInteresting. I run my local llms through ollama and it's zero trouble to get that working in openwebui as long as the ollama server is running.
- diggan 1y agoI think that's the thing. Compared to LM Studio, just running Ollama (fiddling around with terminals) is more complicated than the full E2E of chatting with LM Studio. Of course, for folks used to terminals, daemons and so on it makes sense from the get go, but for others it seemingly doesn't, and it doesn't help that Ollama refuses to communicate what people should understand before trying to use it.
- noman-land 1y agoI love LM Studio. It's a great tool. I'm waiting for another generation of Macbook Pros to do as you did :).
- imranq 1y agoI'd love to host my own LLMs but I keep getting held back from the quality and affordability of Cloud LLMs. Why go local unless there's private data involved?
- mycall 1y agoOffline is another use case.
- seanmcdirmid 1y agoNothing like playing around with LLMs on an airplane without an internet connection.
- asteroidburger 1y agoIf I can afford a seat above economy with room to actually, comfortably work on a laptop, I can afford the couple bucks for wifi for the flight.
- seanmcdirmid 1y agoIf you are assuming that your Hainan airlines flight has wifi that isn't behind the GFW, even outside of cattle class, I have some news for you...
- sach1 1y agoGetting around the GFW is trivially easy.
- seanmcdirmid 1y agoya ya, just buy a VPN, pay the yearly subscription, and then have them disappear the week after you paid. Super trivially frustrating.
- deleted 1y ago[deleted]
- zackify 1y agoI love LM studio but I’d never waste 12k like that. The memory bandwidth is too low trust me. Get the RTX Pro 6000 for 8.5k with double the bandwidth. It will be way better
- marci 1y agoYou can't run deepseek-v3/r1 on the RTX Pro 6000, not to mention the upcomming 1 million context qwen models, or the current qwen3-235b.
- 112233 1y agoI can run full deepseek r1 on m1 max with 64GB of ram. Around 0.5 t/s with small quant. Q4 quant of Maverick (253 GB) runs at 2.3 t/s on it (no GPU offload). Practically, last gen or even ES/QS EPYC or Xeon (with AMX), enough RAM to fill all 8 or 12 channels plus fast storage (4 Gen5 NVMEs are almost 60 GB/s) on paper at least look like cheapest way to run these huge MoE models at hobbyist speeds.
- marci 1y agoIf you're talking about Deepseek r1 with llama.cpp and mmap, then at this point you can run deepseek r1 on a raspberry zero with a 256GB micro sdcard and a phone charger. The only metric left to know is one's patience.
- tymscar 1y agoWhy would they pay 2/3 of the price for something with 1/5 of ram? The whole point of spending that much money for them is to run massive models, like the full R1, which the Pro 6000 cant
- zackify 1y agoBecause waiting forever for initial prompt processing with realistic number of MCP tools enabled on a prompt is going to suck without the most bandwidth possible And you are never going to sit around waiting for anything larger than the 96+gb of ram that the RTX pro has. If you’re using it for background tasks and not coding it’s a different story
- tt726259 1y ago[dead]
- wangbang 1y ago[dead]
- storus 1y agoIf the rumors about splitting CPU/GPU in new Macs are true, your MacStudio will be the last one capable of running DeepSeek R1 671B Q4. It looks like Apple had an accidental winner that will go away with the end of unified RAM.
- phren0logy 1y agoI have not heard this rumor. Source?
- prophesi 1y agoI believe they're talking about the rumors by an Apple supply chain analyst, Ming-Chi Kuo. https://www.techspot.com/news/106159-apple-m5-silicon-rumored-ditch-unified-memory-split.html https://www.techspot.com/news/106159-apple-m5-silicon-rumore...
- diggan 1y agoSeems Apple is waking up to the fact that if it's too easy to run weights locally, there really isn't much sense to having their own remote inference endpoints, so time to stop the party :)
- prophesi 1y agoI thought their goal was to completely remove the need for a remote inference endpoint in the first place? May have read your comment wrong.
- diggan 1y agoNo, I think Apple been clear from the beginning that they won't be able to do everything on the devices themselves, that's why they're building the infrastructure/software for their "cloud intelligence system" or whatever they call it.
- whatevsmate 1y agoI did this a month ago and don't regret it one bit. I had a long laundry list of ML "stuff" I wanted to play with or questions to answer. There's no world in which I'm paying by the request, or token, or whatever, for hacking on fun projects. Keeping an eye on the meter is the opposite of having fun and I have absolutely nowhere I can put a loud, hot GPU (that probably has "gamer" lighting no less) in my fam's small apartment.
- chisleu 1y agoRight on. I also have a laundry list of ML things I want to do starting with fine tuning models. I don't mind paying for models to do things like code. I like to move really fast when I'm coding. But for other things, I just didn't want to spend a week or two coming up on the hardware needed to build a GPU system. You can just order a big GPU box, but it's going to cost you astronomically right now. Building a system with 4-5 PCIE 5.0 x16 slots, enough power, enough pcie lanes... It's a lot to learn. You can't go on PC part picker and just hunt a motherboard with 6 double slots. This is a machine to let me do some things with local models. My first goal is to run some quantized version of the new V3 model and try to use it for coding tasks. I expect it will be slow for sure, but I just want to know what it's capable of.
- datpuz 1y agoI genuinely cannot wrap my head around spending this much money on hardware that is dramatically inferior to hardware that costs half the price. MacOS is not even great anymore, they stopped improving their UX like a decade ago.
- chisleu 1y agoHow can you say something so brave, and so wrong?
- minimaxir 1y agoLM Studio has quickly become the best way to run local LLMs on an Apple Silicon Mac: no offense to vllm/ollama and other terminal-based approaches, but LLMs have many levers for tweaking output and sometimes you need a UI to manage it. Now that LM Studio supports MLX models, it's one of the most efficient too. I'm not bullish on MCP, but at the least this approach gives a good way to experiment with it for free.
- pzo 1y agoI just wish they did some facelifting of UI. Right now is too colorfull for me and many different shades of similar colors. I wish they copy some color pallet from google ai studio or from trae or pycharm.
- nix0n 1y agoLM Studio is quite good on Windows with Nvidia RTX also.
- boredemployee 1y agocare to elaborate? i have rtx 4070 12gb vram + 64gb ram, i wonder what models I can run with it. Anything useful?
- nix0n 1y agoLM Studio's model search is pretty good at showing what models will fit in your VRAM. For my 16gb of VRAM, those models do not include anything that's good at coding, even when I provide the API documents via PDF upload (another thing that LM Studio makes easy). So, not really, but LM Studio at least makes it easier to find that out.
- boredemployee 1y agook, ty for the reply!
- Eupolemos 1y ago
- gregorym 1y ago[flagged]
- simonw 1y agoThat's clearly your own product (it links to Koroworld in the footer and you've posted about that on Hacker News in the past). Are you sharing any of your revenue from that $79 license fee with the https://ollama.com/ https://ollama.com/ project that your app builds on top of?
- cchance 1y agoThe UI's not even as nice as lmstudio lol wtf and their gonna charge 79$?!?!?
- bouke 1y agoIt is even worse; they are offering a commercial product under the same name of the open source project it is based on: https://github.com/kevinhermawan/Ollamac?tab=readme-ov-file#%EF%B8%8F-important-notice https://github.com/kevinhermawan/Ollamac?tab=readme-ov-file#... and https://github.com/gregorym/ollamac-pro/issues/1 https://github.com/gregorym/ollamac-pro/issues/1.
- usef- 1y agoIs this related to the open source ollamac at all? https://github.com/kevinhermawan/Ollamac https://github.com/kevinhermawan/Ollamac
- visiondude 1y agoLMStudio works surprisingly well on M3 Ultra 64gb and 27b models. Nice to have a local option, especially for some prompts.
- squanchingio 1y agoI'll be nice to have the MCP servers exposed like LMStudio OpenAI-like endpoints.
- patates 1y agoWhat models are you using on LM Studio for what task and with how much memory? I have a 48GB macbook pro and Gemma3 (one of the abliterated ones) fits my non-code use case perfectly (generating crime stories which the reader tries to guess the killer). For code, I still call Google to use Gemini.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- robbru 1y agoI've been using the Google Gemma QAT models in 4B, 12B, and 27B with LM Studio with my M1 Max. https://huggingface.co/lmstudio-community/gemma-3-12B-it-qat-GGUF https://huggingface.co/lmstudio-community/gemma-3-12B-it-qat...
- t1amat 1y agoI would recommend Qwen3 30B A3B for you. The MLX 4bit DWQ quants are fantastic.
- redman25 1y agoQwen is great but for creative writing I think Gemma is a good choice. It has better EQ than Qwen IMO.
- deleted 1y ago[deleted]
- api 1y agoI wish LM Studio had a pure daemon mode. It's better than ollama in a lot of ways but I'd rather be able to use BoltAI as the UI, as well as use it from Zed and VSCode and aider. What I like about ollama is that it provides a self-hosted AI provider that can be used by a variety of things. LM Studio has that too, but you have to have the whole big chonky Electron UI running. Its UI is powerful but a lot less nice than e.g. BoltAI for casual use.
- SparkyMcUnicorn 1y agoThere's a "headless" checkbox in settings->developer
- diggan 1y agoStill, you need to install and run the AppImage at least once to enable the "lms" cli which can later be used. Would be nice with a completely GUI-less installation/use method too.
- t1amat 1y agoThe UI is the product. If you just want the engine, use mlx-omni-server (for MLX) or llama-swap (for GGUF) and huggingface-cli (for model downloads).
- diggan 1y agoThose don't offer the same features as LM Studio itself does, even when you don't consider the UI. If there was a "LM Engine" CLI I could install, then yeah, but there isn't, hence the need to run the UI once to get "the engine".
- SparkyMcUnicorn 1y agoI haven't used this and maybe it doesn't solve the problem you're describing, but might be worth looking at. https://github.com/lmstudio-ai/lms https://github.com/lmstudio-ai/lms https://lmstudio.ai/docs/cli https://lmstudio.ai/docs/cli
- b0a04gl 1y ago[dead]
- politelemon 1y agoThe initial experience with LMStudio and MCP doesn't seem to be great, I think their docs could do with a happy path demo for newcomers. Upon installing the first model offered is google/gemma-3-12b - which in fairness is pretty decent compared to others. It's not obvious how to show the right sidebar they're talking about, it's the flask icon which turns into a collapse icon when you click it. I set the MCP up with playwright, asked it to read the top headline from HN and it got stuck on an infinite loop of navigating to Hacker News, but doing nothing with the output. I wanted to try it out with a few other models, but figuring out how to download new models isn't obvious either, it turned out to be the search icon. Anyway other models didn't fare much better either, some outright ignored the tools despite having the capacity for 'tool use'.
- t1amat 1y agoGemma3 models can follow instructions but were not trained to call tools, which is the backbone of MCP support. You would likely have a better experience with models from the Qwen3 family.
- cchance 1y agoThat latter issue isnt a lmstudio issue... its a model issue,
- Thews 1y agoOthers mentioned qwen3, but which works fine with HN stories for me, but the comments still trip it up and it'll start thinking the comments are part of the original question after a while. I also tried the recent deepseek 8b distill, but it was much worse for tool calling than qwen3 8b.
- cylinderthought 1y agogood.
- deleted 1y ago[deleted]
- v3ss0n 1y agoClosed source - wont touch.
- xyc 1y agoGreat to see more local AI tools supporting MCP! Recently I've also added MCP support to recurse.chat. When running locally (LLaMA.cpp and Ollama) it still needs to catch up in terms of tool calling capabilities (for example tool call accuracy / parallel tool calls) compared to the well known providers but it's starting to get pretty usable.
- rshemet 1y agohey! we're building Cactus (https://github.com/cactus-compute https://github.com/cactus-compute), effectively Ollama for smartphones. I'd love to learn more about your MCP implementation. Wanna chat?
- zaps 1y agoNot to be confused with FL Studio
- bbno4 1y agoIs there an app that uses OpenRouter / Claude or something locally but has MCP support?
- eajr 1y agoI've been considering building this. Havent found anything yet.
- cchance 1y agovscode with roocode... just use the chat window :S
- cedws 1y agoI’m looking for something like this too. Msty is my favourite LLM UI (supports remote + local models) but unfortunately has no MCP support. It looks like they’re trying to nudge people into their web SaaS offering which I have no interest in.
- jtreminio 1y agoI’ve been wanting to try LM Studio but I can’t figure out how to use it over local network. My desktop in the living room has the beefy GPU, but I want to use LM Studio from my laptop in bed. Any suggestions?
- skygazer 1y agoUse an openai compatible API client on your laptop, and LM Studio on your server, and point the client to your server. LM Server can serve an LLM on a desired port using the openai style chat completion API. You can also install openwebui on your server and connect to it via a web browser, and configure it to use the LM Studio connection for its LLM.
- numpad0 1y ago[>_] -> [.* Settings] -> Serve on local network ( o) Any OpenAI-compatible client app should work - use IP address of host machine as API server address. API key can be bogus or blank.
- sixhobbits 1y agoMCP terminology is already super confusing, but this seems to just introduce "MCP Host" randomly in a way that makes no sense to me at all. > "MCP Host": applications (like LM Studio or Claude Desktop) that can connect to MCP servers, and make their resources available to models. I think everyone else is calling this an "MCP Client", so I'm not sure why they would want to call themselves a host - makes it sound like they are hosting MCP servers (definitely something that people are doing, even though often the server is run on the same machine as the client), when in fact they are just a client? Or am I confused?
- guywhocodes 1y agoMCP Host is terminology from the spec. It's the software that makes llm calls, build prompts, interprets tool call requests and performs them etc.
- sixhobbits 1y agoSo it is, I stand corrected. I googled mcp host and the lmstudio link was the first result. Some more discussion on the confusion here https://github.com/modelcontextprotocol/modelcontextprotocol/discussions/135 https://github.com/modelcontextprotocol/modelcontextprotocol... where they acknowledge that most people call it a client and that that's ok unless the distinction is important. I think host is a bad term for it though as it makes more intuitive sense for the host to host the server and the client to connect to it, especially for remote MCP servers which are probably going to become the default way of using them.
- kreetx 1y agoI'm with you on the confusion, it makes no sense at all to call it a host. MCP host should host the MCP server (yes, I know - that is yet a separate term). The MCP standard seems a mess, e.g take this paragraph from here[1] > In the Streamable HTTP transport, the server operates as an independent process that can handle multiple client connections. Yes, obviously, that is what servers do. Also, what is "Streamable HTTP"? Comet, HTTP2, or even websockets? SSE could be a candidate, but it isn't as it says "Streamable HTTP" replaces SSE. > This transport uses HTTP POST and GET requests. Guys, POST and GET are verbs for HTTP protocol, TCP is the transport. I guess they could say that they use HTTP protocol, which only uses POST and GET verbs (if that is the case). > Server can optionally make use of Server-Sent Events (SSE) to stream multiple server messages. This would make sense if there weren't the note "This replaces the HTTP+SSE transport" right below the title. > This permits basic MCP servers, as well as more feature-rich servers supporting streaming and server-to-client notifications and requests. Again, how is streaming implemented (what is "Streaming HTTP")?. Also, "server-to-client .. requests"? SSE is unidirectional, so those requests are happening over secondary HTTP requests? -- And then the 2.0.1 Security Warning seems like a blob of words on security, no reference to maybe same-origin. Also, "for local servers bind to localhost and then implement proper authentication" - are both of those together ever required? Is it worth it to even say that servers should implement proper authentication? Anyway, reading the entire documentation one might be able to put a charitable version of the MCP puzzle together that might actually make sense. But it does seem that it isn't written by engineers, in which case I don't understand why or to whom is this written for. [1] https://modelcontextprotocol.io/specification/draft/basic/transports#streamable-http https://modelcontextprotocol.io/specification/draft/basic/tr...
- mkagenius 1y agoOn M1/M2/M3 Mac, you can use Apple Containers to automate[1] the execution of the generated code. I have one running locally with this config: { "mcpServers": { "coderunner": { "url": "http://coderunner.local:8222/sse" } } } 1. CodeRunner: https://github.com/BandarLabs/coderunner https://github.com/BandarLabs/coderunner (I am one of the authors)
- smcleod 1y agoI really like LM Studio but their license / terms of use are very hostile. You're in breach if you use it for anything work related - so just be careful folks!
- jmetrikat 1y agogreat! it's very convenient to try mcp servers with local models that way. just added the `Add to LM Studio` button to the anytype mcp server, looks nice: https://github.com/anyproto/anytype-mcp https://github.com/anyproto/anytype-mcp
- elizabethadavis 1y ago[dead]
- b0dhimind 1y agoI wonder how LM Studio and AnythingLLM contrasts especially in upcoming months... I like AnythingLLM's workflow editor. I'd like something to grow into for my doc-heavy job. Don't want to be installing and trying both.