7 ms·
Ollama Web Search
- chungus42 1y agoMy biggest gripe with small models has been the inability to keep it informed with new data. Seems like this at least eases the process.
- mchiang 1y agoI was pleasantly surprised on the model improvements when testing this feature. For smaller models, it can augment it with the latest data by fetching it from the web, solving the problem of smaller models lacking specific knowledge. For larger models, it can start functioning as deep research.
- tripplyons 1y agoJust set up SearXNG locally if you want a free/local web search MCP: https://gist.github.com/tripplyons/a2f9d8bd553802f9296a7ec3b46760de https://gist.github.com/tripplyons/a2f9d8bd553802f9296a7ec3b...
- mchiang 1y agoI haven't tried SearXNG personally. How does it compare to Ollama's web search in terms of the search content returned?
- tripplyons 1y agoI have no idea how well Ollama's works, but I haven't ran into any issues with SearXNG. The alternatives aren't worth paying for in any use case I've encountered.
- deleted 1y ago[deleted]
- disiplus 1y agoThat's what i have together with open webui and gpt-oss-120b. it works reasonably well. But sometimes the searches are slow.
- tripplyons 1y agoYou can try removing search engines that fail or reducing their timeout setting to something faster than the default of a few seconds.
- disiplus 1y agoSearXNG is fast, its mostly the code that triggers the searches. Because, my daily is chatgpt, i still did not try to tweak it.
- tripplyons 1y agoI haven't needed to tweak mine for similar reasons, but I'm surprised to hear that the "code that triggers the searches" is slow. Are you referring to something in Open WebUI?
- disiplus 1y agoIt's tools that you can install from open webui https://openwebui.com/tools https://openwebui.com/tools
- sorenjan 1y agoI had no idea they had their own cloud offering, I thought the whole point of Ollama was local models? Why would I pay $20/month to use small inferior models instead of using one of the usual AI companies like OpenAI or even Mistral? I'm not going to make an account to use models on my own computer.
- mchiang 1y agoFair question. Some of the supported models are large and wouldn't fit on most local devices. This is just the beginning, and Ollama does not need to exclude cloud hosted frontier models either with the relationship we've built with the model providers. We just have to be mindful and understand that Ollama stands with developers, and solve the needs. https://ollama.com/cloud https://ollama.com/cloud
- sorenjan 1y ago> Some of the supported models are large and wouldn't fit on most local devices. Why would I use those models on your cloud instead of using Google's or Anthropic's models? I'm glad there are open models available and that they get better and better, but if I'm paying money to use a cloud API I might as well use the best commercial models, I think they will remain much better than the open alternatives for quite some time.
- mchiang 1y agoWhen we started Ollama, we were told how open-source (open-weight wasn't a term back then) will always be inferior to the close-sourced models. This was 2 years ago (Ollama's birthday is July 18th, 2023). Fast forward to now, open models are quickly catching up, and at a significantly lower price point for most and can be customized for specific tasks instead of being general purpose. For general purpose models, absolutely the closed models are currently dominating.
- typpilol 1y agoYa a lot of ppl don't realize you could spend 2k on a 5090 to run some of the large models. Or spend 20 a month for models even a 5090 couldn't run. And not have to spend your own electricity, hardware, maintenance, updates etc.
- throwaway12345t 1y agoDo they pull their own index like brave or are they using Bing/Google in the background?
- tripplyons 1y agoBased on the fact that there are very few up-to-date English-language search indexes (Google, Bing, and Brave if you count it), it must be incredibly costly. I doubt they are maintaining their own.
- throwaway12345t 1y agoWe need more indexes
- JumpCrisscross 1y ago> We need more indexes Not particularly. Indexes are sort of like railroads. They're costly to build and maintain. They have significant external costs. (For railroads, in land use. For indexes, in crawler pressure on hosting costs.) If you build an index, you should be entitled to a return on your investment. But you should also be required to share that investment with others (at a cost to them, of course).
- ineedasername 1y agoDo we know what OpenAI uses? Have they built their own, or piggy back on moneybags $MS and Bing?
- tripplyons 1y agoThey use Bing: https://www.forbes.com/sites/katherinehamilton/2023/05/23/chatgpt-will-now-have-access-to-real-time-info-from-bing-search-report-says/ https://www.forbes.com/sites/katherinehamilton/2023/05/23/ch...
- tripplyons 1y agoMore competition in the space would be great for me as a consumer, but the problem is that the high fixed costs make starting an index difficult.
- simonw 1y agoI'd love to know what search engine provider they're using under the hood for this. I asked them on Twitter and didn't get a reply (yet) https://twitter.com/simonw/status/1971210260015919488 https://twitter.com/simonw/status/1971210260015919488 Crucially, I want to understand the license that applies to the search results. Can I store them, can I re-publish them? Different providers have different rules about this.
- mchiang 1y agoWe work with search providers and ensure that we have zero data retention policies in place. The search results are yours to own and use. You are free to do what you want with it. Of course you are bound by local laws of the legal jurisdiction you are in.
- simonw 1y agoOK, so it looks like you aren't willing to share which providers you are working with. Can you share the rationale for not sharing that information instead?
- mchiang 1y agoWe have relationships with many providers and I don't want to be seen as promoting or not promoting a specific provider. Some decent privacy-preserving vendors - Brave, Exa, Parallel Web Systems, DuckDuckGo etc We will continue to monitor what's good to improve the output quality and results. Sometimes it could be the combination of providers to yield even better results. If I say one combination right now, and realize another combination is better, and make changes, I wouldn't need to broadcast it each time or risk misrepresenting the feature, which is to have amazing search and research capabilities that can augment models for a superior output.
- dcreater 1y agoThis information is very useful to the open source community. Whats the rationale in not "building in the public"? Is Ollama turning its back on the open source community? Also why should we believe ollama web search is better than my locally run searxng server?
- MisterBiggs 1y agoI was hoping for more details about their implementation, I saw ollama as the open source // platform agnostic tool but I worry their recent posturing is going against that
- jmorgan 1y agoWe did consider building functionality into Ollama that would go fetch search results and website contents using a headless browser or similar. However we had a lot of worries about result quality and also IP blocking from Ollama creating crawler-like behavior. Having a hosted API felt like a fast path to get results into users' context window, but we are still exploring the local option. Ideally you'd be able to stay fully local if you want to (even when using capabilities like search)
- dcreater 1y agoTheir posture has continually been getting worse and worse. It's deceptive and I've expunged it from all my systems
- wirybeige 1y agoTheir GUI is closed-source. If someone wants an easy to use & easy to setup app, may as well use LMStudio, which doesn't try to pretend to be OSS. Or use ramalama which is basically just containerizing LLMs and the relevant bits, pretty damn similar to ollama. Or just go back to "basics" and use llama.cpp or vllm.
- bigyabai 1y ago> Create an API key from your Ollama account. Dead on arrival. Thanks for playing, Ollama, but you've already done the leg work in obsoleting yourself.
- disiplus 1y agothey had at some point start earning money.
- bigyabai 1y agoAt some point you have to earn user trust. If Ollama won't be the Open Source Ollama API provider, there are several endpoint-compatible alternatives happy to replace them. From where I'm standing, there's not enough money in B2C GPU hosting to make this sort of thing worthwhile. Features like paid search APIs this really hammer home how difficult it is to provide value around that proposition.
- timothymwiti 1y agoDoes anyone know if the python and JavaScript examples on the blog work without an Ollama Account?
- mrkeen 1y agoAny tips on local/enterprise search? I like using ollama locally and I also index and query locally. I would love to know how to hook ollama up to a traditional full-text-search system rather than learning how to 'fine tune' or convert my documents into embeddings or whatnot.
- ineedasername 1y agoYou can use solr, very good full text search and it has an mcp integration. That’s sufficient on its own and straightforward to setup: https://github.com/mjochum64/mcp-solr-search https://github.com/mjochum64/mcp-solr-search A slightly heavier lift, but only slightly, would be to also use solr to also store a vectorized version of your docs and simultaneously do vector similarity search, solr has built in knn support fort it. Pretty good combo to get good quality with both semantic and full-text search. Though I’m not sure if it would be relatively similar work to do solr w/ chromadb, for the vector portion, and marry the result stewards via llm pixie dust (“you are the helpful officiator of a semantic full-text matrimonial ceremony” etc). Also not sure the relative strengths of chromadb vs solr on that- maybe scales better for larger vector stores?
- all2 1y agodocling might be a good way to go here. Or consider one of the existing full text search engines like Typesense.
- lxgr 1y agoDoes this work with (tool use capable) models hosted locally?
- yggdrasil_ai 1y agoI don't think ollama officially supports any proper tool use via api.
- lxgr 1y agoHuh, I was pretty sure I used it before, but maybe I’m confusing it with some other python-llm backend. Is https://ollama.com/blog/tool-support https://ollama.com/blog/tool-support not it?
- all2 1y agoIt depends on the model. Deepseek-R1 says it supports tool use, but the system prompt template does not have the tool-include callouts. YMMV
- parthsareen 1y agoHi - author of the post. Yes it does! The "build a search agent" example can be used with a local model. I'd recommend trying qwen3 or gpt-oss
- lxgr 1y agoVery cool, thank you! Looking forward to try it with a few shell scripts (via the llm-ollama extension for the amazing Python ‘llm’) or Raycast (the lack of web search support for Ollama has been one of my biggest reasons for preferring cloud-hosted models).
- parthsareen 1y agoSince we shipped web search with gpt-oss in the Ollama app I've personally been using that a lot more especially for research heavy tasks that I can shoot off. Plus with a 5090 or the new macs it's super fast.
- yggdrasil_ai 1y agoI wish they would instead focus on local tool use. I could just use my own web search via brave api.
- parthsareen 1y agoHey! Author of the blogpost and I also work on Ollama's tool calling. There has been a big push on tool calling over the last year to improve the parsing. What's the issues you're running into with local tool use? What models are you using?
- vrzucchini 1y agoHey, unrelated to the question you're answering but where do I see the rate limits for free and paid tiers?
- yggdrasil_ai 1y agoI went back and had another look at my implementation, and got it to work. Sorry I was mistaken!
- coffeecoders 1y agoOn a slightly related note- I've been thinking about building a home-local "mini-Google" that indexes maybe 1,000 websites. In practice, I rarely need more than a handful of sites for my searches, so it seems like overkill to rely on full-scale search engines for my use case. My rough idea for architecture: - Crawler: A lightweight scraper that visits each site periodically. - Indexer: Convert pages into text and create an inverted index for fast keyword search. Could use something like Whoosh. - Storage: Store raw HTML and text locally, maybe compress older snapshots. - Search Layer: Simple query parser to score results by relevance, maybe using TF-IDF or embeddings. I would do periodic updates and build a small web UI to browse. Anyone tried it or are there similar projects?
- fabiensanglard 1y agoHave you ever tried https://marginalia-search.com https://marginalia-search.com ? I love it.
- matsz 1y agoYou could take a look at the leaked Yandex source code from a few years ago. I'd believe their architecture should be decent enough.
- harias 1y ago
- frabonacci 1y agoThis is a nice first step - web search makes sense, and it’s easy to imagine other tools being added next: filesystem, browser, maybe even full desktop control. Could turn Ollama into more than just a model runner. Curious if they’ll open up a broader tool API for third-party stuff too
- drnick1 1y agoWhat "Ollama account?" I am confused, I thought the point of Ollama was to self-host models.
- mchiang 1y agoTo provide additional features or using Ollama's cloud hosted models, you can signup for an Ollama account. For starter, this is completely optional. It can be completely local too for you to publish your own models to ollama.com that you can share with others.
- dumbmrblah 1y agoWhat is the data retention policy for the free account versus the cloud account?
- nextworddev 1y agoCan someone tell me how much this costs and how this compares to Tavily etc
- typpilol 1y agoTaviy gives you 1k free requests a month. Even with heavy ai usage I'm only at like 400/1000 for the month
- anonyonoor 1y agoI know it might be a security nightmare, but I still want to see an implementation of client-side web search. Like a full search engine that can visit pages on your behalf. Is anyone building this?
- not_really 1y agosounds like a good way to get your IP flagged by cloudflare
- apimade 1y agoAgenticSeek, or you can get pretty far with local qwen and Playwright-Stealth or SeleniumBase integrated directly into your Chrome (running with Chrome DevTools Protocol enabled).
- riskable 1y agoWTF is going to happen to Google's ad revenue if every PC has an AI that can perform searches on the user's behalf?
- tartoran 1y agoThey'll have to squeeze it all from Youtube!
- andrewmcwatters 1y agogoogle.com/sorry
- thimabi 1y agoThey can always pivot to their Search-via-API business :) It takes lots of servers to build a search engine index, and there’s nothing to indicate that this will change in the near future.
- onesociety2022 1y agoHow is that any different than someone installing an ad blocker in their browser? Arguably ad blocker is much simpler technology than running a local LLM and has been available for years now. And yet Google’s ad revenue seems to have remained unaffected.
- hadlock 1y agoIt's been demonstrated that as ChatGPT usage goes up, traffic to sites dependent on SEO search ranking has gone down, roughly proportionally, every month over the last ~18 months. ChatGPT is free and fast and requires no technical know-how. Installing an ad blocker requires knowing what one is, and the time and energy to install a browser plugin. Pretty much everyone I know thinks free online ChatGPT type products is an absolute existential thread to Google's ad dominance. Even mediocre LLMs provide a vastly better experience than ad choked pages linking to ad choked SEO optimized websites serving (largely) google's own ads.
- system2 1y agoThere are millions of websites, and a local LLM cannot scrape all of them to make sense of them. Think about it. OpenAI can do it because they spend millions to train its systems. Many sites have hidden sitemaps that cannot be found unless submitted to google directly. (Not even listed in robots txt most of the time). There is no way a local LLM can keep up with up to date internet.
- kordlessagain 1y agoI have a MCP tool that uses SERP API and it works quite well.
- andrewmutz 1y agoOllama is a business? They raised money? I thought it was just a useful open source product. I wonder how they plan to monetize their users. Doesn't sound promising.
- coolspot 1y agoThey are former Docker employees running Docker playbook.
- Cheer2171 1y ago[flagged]
- cristoperb 1y agoOllama is a ycombinator startup, so I guess they have to find some roi at some point.[1] I personally found Ollama to be an easy way to try out local LLMs and appreciate them for that (and I still use it to download small models on my laptop and phone (via termux)), but I've long switched to llama.cpp + llama-swap[2] on my dev desktop. I download whatever ggufs I want from hugging face and just do `git pull` and `cmake --build build --config Release` from my llama.cpp directory whenever I want to update. 1: https://www.ycombinator.com/companies/ollama https://www.ycombinator.com/companies/ollama 2: https://github.com/mostlygeek/llama-swap https://github.com/mostlygeek/llama-swap
- blihp 1y agoThere are very few recently launched pure open source projects these days (most are at least running donation-ware models or funded by corporate backers), none in the AI space that I'm aware of.
- brabel 1y agoWell the real open source project is llama.cpp which Ollama basically wrapped and made a nice interface on top of. Now they do more things as they want to be a real business, but llama.cpp is now doing most things people wanted from something like ollama, like serving a REST API compatible with OpenAPI, downloading and managing local LLMs… while remaining an actual open source project without VC money as far as I know.
- thomastraum 1y agoI am just working on a tool using websearch and iterating over different providers. openAI, xAI, gemini all suffer from not being allowed on respective competitor sites. this searched works for me with some quick tests well on YT videos, which OpenAI web search can't access. It kind of failed on X but sometimes returned ok relevant results. Definitely hit and miss but on average good
- chrisshroba 1y agoAre the rate limits documented somewhere?
- andai 1y agoI added search to my LLMs years ago with the python DuckDuckGo package. However I found that Google gives better results, so I switched to that. (I forget exactly but I had to set up something in a Google dev console for that.) I think the DDG one is unofficial, and the Google one has limits (so it probably wouldn't work well for deep research type stuff). I mostly just pipe it into LLM apis. I found that "shove the first few Google results into GPT, followed by my question" gave me very good results most of the time. It of course also works with Ollama, but I don't have a very good GPU, so it gets really slow for me on long contexts.
- ivape 1y agoHow do you meaningfully use it without using scraping APIs? Aren't the official apis severely limited?
- selcuka 1y agoGoogle Programmable Search Engine [1] is pretty good if your needs are within their usage limits. [1] https://programmablesearchengine.google.com/about/ https://programmablesearchengine.google.com/about/
- andai 1y agoThat's the one I use, yeah! You set it up here: https://programmablesearchengine.google.com/controlpanel/create https://programmablesearchengine.google.com/controlpanel/cre... And then it's just a GET: import os import json from req import get url = "https://customsearch.googleapis.com/customsearch/v1" def search(query): data = { "q": query, "cx": os.getenv('GOOGLE_SEARCH_API_KEY'), "key": os.getenv('GOOGLE_SEARCH_API_ID') } results_json = get(url, data) results = json.loads(results_json) results = results["items"] return results
- Cheer2171 1y agoYour regular reminder that you don't need ollama to get a quick chat engine on the command line, you can just do this with pretty much any major model on huggingface: pip install transformers transformers chat Qwen/Qwen2.5-0.5B-Instruct
- mmaunder 1y agoSo, use ollama to avoid cloud models and services, but ollama sells cloud models and services. The dissonance makes my teeth hurt.
- orliesaurus 1y agoExa, Tavily or Firecrawl. Which one is it?
- jerrygoyal 1y agoI'm looking to use web search in production, but they haven't mentioned the price. Only thing that's mentioned is $20/month, but how much quota does it include?
- mchiang 1y agoSorry about this. We are working really hard on providing a usage based pricing. During the preview period we want to start offering a $20 / month plan tailored for individuals - and we are monitoring the usage and making changes as people hit rate limits so we can satisfy most use cases, and be generous.
- enoch2090 1y agoThat's the essence of these services, they never explicitly mention the quota, or secretly lowers it at some point.
- alberth 1y agoDumb question: is this affiliated with Meta? Or is this just someone trying to monetize Meta open source models?
- mchiang 1y agoNo, Ollama is it's own project and separate. You can check it out via GitHub https://github.com/ollama/ollama https://github.com/ollama/ollama
- mostMoralPoster 1y ago[dead]
- kgeist 1y agoI use Llama.cpp with Tavily search (they give free credits each month). LibreChat has built-in support for it. No Ollama needed.
- Tepix 1y agoLooks like Ollama is focusing more and more on non-local offerings. Also their performance is worse than say vLLM. What's a good Ollama alternative (for keeping 1-5x RTX 3090 busy) if you want to run things like open-webui (via an OpenAI compatible API) where your users can choose between a few LLMs?
- Ey7NFZ3P0nzAe 1y agoi heard about Llamaswap and vllm
- kgeist 1y agoAt work I've set up LibreChat + LlamaSwap + llama.cpp 200 weekly users :)
- Tepix 1y agoHow do you deal with different users wanting to use different LLMs at the same time?
- tempodox 1y agoIs the web search also integrated into the locally running native ollama binaries, and if so, how can I use it?
- Boristoledano 1y ago[dead]