6 ms·
Turning a pile of documents into a searchable useable knowledge base
- linuxrebe1 2mo agoI had an issue. A documents folder with over 12k objects in it. A hodgepodge of folders and sub-folders. That over time had created a mess that no amount of file movement was ever going to make it usable. I wanted: 1) To keep my data local 2) be able to filter out PII and other data 3) Be able to find and delete duplicates 4) Get short synopsis of what a document is 5) Semantic and keyword search 6) All of this kept local to me requiring no internet access and no tokens spent to train someone elses AI. The result I call DocuBrowser and in it's current form is FOSS (GPL-3) licensed for your personal use. The UI is in your browser. The AI models used are held local and are tiny, Available for Linux(RPM,Deb, and tgz) Windows and Mac. Let me know what you think and thanks for taking the time to try it out.
- bobim 2mo agoCould it be extended so it also extracts pictures from pptx and xlsx and run vision to get a description to be added to the text content before indexing?
- linuxrebe1 2mo agoLet me look into this
- clif_mcIrvin 2mo agoHow about jpegs or other scanner images files? We have hundreds of scanned documents that were never pdf wrapped.
- linuxrebe1 2mo agohmmmm :)
- esperent 2mo agoI've been working on something related - extracting tons of data from various formats to allow searching them - and the solution I chose for xlxs and xls files was headless LibreOffice to convert them to CSV. There's also exceljs but I found it didn't work for many old xls files. I didn't find screenshotting of spreadsheets worked well, vision wasn't very accurate on them. I do use it for PDFs though. For docx it's probably fine either way but I went with LibreOffice -> markdown.
- linuxrebe1 2mo agoI went with the python libraries (pydoc and pyxls for example), because it's portable and doesn't require a big download to a users system if they don't already have it installed.
- bobim 2mo agoMy take was on pictures embedded into those documents, I'm not sure screenshotting would help as the text/numeric data is already there. Just saying.
- seb1204 2mo agoSounds similar to https://docs.paperless-ngx.com/ https://docs.paperless-ngx.com/ Key difference I see is that you point it to a folder instead of uploading to a system.
- vsviridov 2mo agoI think paperless devs are working on AI integration, and there are 3rd party solutions. I'm holding out for an official one, so far. It's pretty cool, I've set up a share where the scanner scans, and it automatically picks it up from there and ingests it into the system.
- password4321 2mo agoPersonal use? I need this at work, dragging useful info from tarpits like Teams and GitLab. Also need to search git repos including all branches and history (TIL/xkcd#153'd GitLab's web search can basically only do one branch at a time).
- linuxrebe1 2mo agoI creating DocuRepo as well. though not as fleshed out.
- rukshn 2mo agoBut how’d you access teams when it’s work teams and don’t have api access ?
- password4321 2mo agoMicrosoft Graph API
- gatnoodle 2mo agoThis looks really cool. Can you tell me the minimum specs required to run this? It would nice if you could add it to the readme as well.
- linuxrebe1 2mo agoI've run it on a VM with 4G ram and no GPU. It runs, But I really recommend 8G ram at least. If you have a GPU (like I do) with 4G vRAM that is ideal. Will get this in the readme. Thanks for the suggestion. I really tried to build this to minimal spec.
- fnordian 2mo agoIt’s either restricted to personal use, or it’s GPL-3. How can you have both?
- linuxrebe1 2mo agoBy restricted for personal use I mean it's not networked. It's running on your system only. It's not a networked commercial product able to do SSO etc. It's not an enterprise level product.
- aucisson_masque 2mo agoI'm a huge fan of recall, going to test this out. This looks very interesting.
- rahimnathwani 2mo agoDid you mean Recoll (https://www.recoll.org/ https://www.recoll.org/)?
- aucisson_masque 2mo agoIndeed
- asciimoo 2mo agoWe need projects like this. Automatically classifying the files is smart. I'm working on a similar application called Hister (https://github.com/asciimoo/hister https://github.com/asciimoo/hister). I should borrow some of your ideas. =]
- hankbond 2mo agoI have not set up Hister yet but it's on my list to try out. How would I do something like host it on my Unraid box but have it index/persist my local MacBook browsing history?
- linuxrebe1 2mo agoI just had a wild thought. Combine Hister with my RepoSearch app. Point it at a companies Internal github/gitlab and have a searchable knowledge base of your git repos.
- asciimoo 2mo agoI like the idea. Could you share your RepoSearch app? Also, we have Discord & IRC channels, please join and start brainstorming.
- NKosmatos 2mo agoLooks good, definitely going to try it. Extra thanks for creating something fully local, we need more projects like this one!
- linuxrebe1 2mo agothankyou
- toomuchtodo 2mo agoHow do you feel about supporting an S3 compatible target as a feature request?
- linuxrebe1 2mo agoI'm actually thinking of this for a commercial product feature. However, if you use a tool like Rclone on Windows, Linux or Mac. Mount the s3 bucket and you can then run DocuBrowse as if the s3 bucket were local.
- subhobroto 2mo agoI love your project on many fronts. One, you're using Claude. Two, you used Python - but most importantly, you personally care about it. I will be using this, and I will be making contributions to it as well. > I'm actually thinking of this for a commercial product feature Would you consider writing down which features you would like to make commercial product features and how you would like to price them?
- linuxrebe1 2mo agoConsider it yes, However having experience in this ... not really. For now there is a file called Decisions.md in the repo that is my "notes to self" if you will about where and what I need to do.
- deleted 2mo ago[deleted]
- drizzler 2mo agoI just installed this and, after a few hiccups, got it up and running on my Ubuntu system. Works great, looks great. Thank you for this. Half of my documents are OpenDocument format. Is there any chance you'll be supporting ODF in the future?
- linuxrebe1 2mo agoYes, not supporting it is an oversight I will correct.
- linuxrebe1 2mo agoWill have version 0.9.1 out later today to support ODF formats.
- linuxrebe1 2mo agov0.9.1 is in the repo and packages have been built. It now does all of the ODT formats.
- jphorism 2mo agoNice, what are you hoping to accomplish with this project?
- NamlchakKhandro 2mo agoA resume
- passwordoops 2mo agoCare to elaborate?
- linuxrebe1 2mo ago- Filling a need I personally have. - Learning how to leverage AI for real world use not just to fill up a data center. - Personal knowledge -developing skills Pretty much in that order
- Avery29 2mo agoThe hardest part of these projects is usually not making documents searchable
- karmakaze 2mo agoI learned a solution is to turn the documents into vectors in say PostgreSQL (with pgvector) and do a cosine similarity search with a search vector. Doing a search for embed models on HuggingFace shows nomic-ai/nomic-embed-text-v1.5 and Qwen/Qwen3-Embedding-0.6B. I might have used a larger one like Qwen/Qwen3-Embedding-4B. There's some info for AnythingLLM[0] which supports RAG. AnythingLLM has LanceDB out of the box but also supports others including pgvector. [0] https://docs.anythingllm.com/features/embedding-models https://docs.anythingllm.com/features/embedding-models
- mune2gu-chan 2mo agoNot a fan of pushing every personal document to someone else's cloud. Nice to see a tool that keeps everything on disk instead.
- nickweb 2mo agoHonestly. This with Paperless-NGX might be game changing if both pointed to the same folder.
- kamranjon 2mo agoWanted to share Antfly which I think serves a similar niche: https://antfly.io/ https://antfly.io/ https://github.com/antflydb/antfly https://github.com/antflydb/antfly They’ve put a lot of effort into optimizing the local llm pipelines and I have a lot of faith in the devs working on it.
- appstorelottery 2mo agoAnyone getting a bunch of permission errors when running (e.g. Traceback (most recent call last): File "/Users/tron/Applications/DocuBrowse/venv/lib/python3.9/site-packages/psutil/_psosx.py", line 347, in wrapper return fun(self, args, *kwargs) File "/Users/tron/Applications/DocuBrowse/venv/lib/python3.9/site-packages/psutil/_psosx.py", line 508, in net_connections rawlist = cext.proc_net_connections(self.pid, families, types) PermissionError: [Errno 1] Operation not permitted (originated from proc_pidinfo(PROC_PIDLISTFDS) 1/2)
- appstorelottery 2mo agoLiving in bizarro world of AI. Install open source project, fails, feed into OpenCode w/DeepSeekFlash 4 -> feed error into it get fixed. The kill_port function only catches ImportError from the psutil block, so when psutil is installed but raises AccessDenied (common on macOS), it crashes instead of falling back to lsof. In platform_paths.py - add two lines after line 250: except psutil.Error: pass Fixed. Now when psutil raises AccessDenied (as happens on macOS without elevated privileges), it falls through to the lsof/fuser fallback instead of crashing. Try docubrowser start again.
- appstorelottery 2mo agoDisappointed that it wasn't returning a list of paragraphs from eBooks that semantically match; search only appears to list the publications - not the actual match within the document.
- linuxrebe1 2mo agoNoted the bug.
- linuxrebe1 2mo agoI fixed this in version 0.9.1 (just released) thanks for the bug (seriously)
- linuxrebe1 2mo ago
- Ozzie_osman 2mo agoThis is really cool. Can it play nice with gdrive or Dropbox? For better or worse, that's just where my data lives now but I'd love this layer.
- linuxrebe1 2mo agoI use rclone to "mount" them locally. Then it becomes searchable.
- hunmernop 2mo agoCan you make a dockerfile and docker compose file?
- LawrenceKerr 2mo agoMany such open source projects already (which is fantastic). I lose track of them. Today it happened I needed a simple way to embed & query 1TB+ of documents, and I was looking at open source options. Can anyone tell me what their go-to solution is now? Could this be the one? And what are the key differences vs. other open source RAG tools like kotaemon?