5 ms·
Knowledge Should Not Be Gated
- MelonUsk 3mo agoYep, knowledge should not be gated: Imagine Google search without any links or sources named This is the “modern” AI chatbot: It never mentions the training data it used, in fact has no idea what it used (often FB, Reddit and partisan websites) Update: I added the reply about after the fact Googling chatbots do - it’s different
- dmortin 3mo agoSecifically in Google AI Overview I always see links to sites where the information is sourced from. Or at least some of the sites, if the same info is sourced from 100 pages then it only shows 2 or 3, maybe the ones with the biggest PageRanks.
- MelonUsk 3mo agoYep, that’s true But those links are Googled after the model started to answer, they are not the links to the training data Imagine an artificial “librarian” that read all the books and spits hallucinated quotes for you But doesn’t let you enter the library, open a single book or even see the sources for those hallucinated quotes But instead Googles some sources based on hallucinations after generating them ;-) It’s better than nothing but you can Google them, too, while training data (the library) is completely hidden from you, even the public domain parts of it - zero attribution
- dmortin 3mo agoThere should be at least some correlation. When building the model they give more weight to some pages (e.g. Wikipedia) which have bigger trust (pagerank?). And when they provide links in answers, those matches are listed first which have better pagerank for the query. So if it sources something in Wikipedia, it is more likely to provide Wikipedia as a trusted source for it. The problem is when an answer is hallucinated, false, it may provide a source for it which contains the invalid info.
- MelonUsk 3mo agoYep, a few non-profits work on direct training data attribution: OlmoTrace, Guide Labs with Clarity and a few more Labs train the model with attribution baked-in and they say the bigger the model - the more interpretable it becomes Pretty sure it’s the future
- nephihaha 3mo agoSadly it has been during most of human history. I think the establishment resents the masses becoming over educated. The 1990s internet had a wealth of views and information on it. Now you can only access approved sources via search engines thanks to scaremongering, and have CloudFlare monitoring everything you do.
- rightbyte 3mo agoIt seems beyond naive, rather malicious, to upload any useful private data to SaaS LLMs. Like, you are letting them data mine your business. Why are corporations not panicing over this?
- deleted 3mo ago[deleted]
- m11a 3mo agoMost corporations likely have zero data retention agreements with LLM providers, at least for API usage. (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)
- dataflow 3mo ago> (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.) Is actual ZDR verbiage in contracts more specific and limited in scope than what we see advertised publicly ("...except where needed to comply with law or combat misuse" in Anthropic's case)? Because those seem pretty damn vague and large enough holes to drive trucks through.
- lukewarm707 3mo agoto combat misuse, we must store and read all prompts and responses. ;) to comply with the law, we must send to the police our detections of illegal activity >:| a guy subpeonaed your chats, i guess we stored them (oops) so now it's illegal to destroy it...
- m11a 3mo agoIt depends on the model provider. OpenAI's is very limited and precisely written. Plus, open-source models hosted on SaaS inference providers tend to come with a strong ZDR agreement too.
- drunken_thor 3mo agoSdks/libs, especially open source sdks, were never about gated knowledge. They were about the providing company making it as easy as possible for you to integrate. You would not need to know the idiosyncrasies behind api retries, paging, rate limits, auth flow, and on and on. The third party developers needed a resource, they call a method and get it. Open source libraries especially are about pooling knowledge, not gating it. This is propaganda for pooling that knowledge inside a service you have to pay to use, and instead of developers all using and improving the same codebase together, they have to spend money to rewrite the same code repeatedly. This is AI companies further trying to undercut open source because it’s free.
- internet2000 3mo agoInformation wants to be free! I remember when that was the rallying cry of hackers. I miss those days.
- sghiassy 3mo ago“Hack the Planet!” — Hackers
- toofy 3mo agothose days are still here, people just want to know where the information is coming from. the rallying cry from hackers has never ever been “information wants to be free from sources” and hackers have also never implied “information from a dipshit should hold the exact same weight as an expert” yet somehow both of those is the world we’re running towards as fast as we can.
- 5701652400 3mo agonow that any software/knowledge is copyable given sufficient cash and AIs, gating knowledge migth be the only thing that protects your business. otherwise you do not have business.
- bonoboTP 3mo agoAI generated article.
- Diti 3mo agoYep. Also no author – an actual writer would have taken credit. I flagged the article but it did nothing.
- dofm 3mo agoWe are at the breathless-but-low-information-posts-about-plain-text-formats point in the cycle.
- Herring 3mo agoThis comment is another example of "Nobody ever gets credit for fixing problems that never happened." There's a massive push to add unnecessary complexity to everything out there, because complexity pays all our bills.
- dofm 3mo agoOh I don't disagree. I'm not saying it's wrong. It's always right at some point in the cycle. I'm just saying it is cyclical. Databases => plain text => single-file databases, repeat. shared hosting => dedicated hosting => vms => jamstack, repeat, etc. Can't sell complexity without oversimplicity or simplicity without overcomplexity. But this is quite a long blog post, with typical blog flourishes, about not very much.
- jdw64 3mo agoPersonally, I think the ability to distinguish between all the knowledge that's overflowing is becoming a characteristic of the current establishment. In reality, the number of sites where you can get good information is extremely limited. It feels like we're in an era where discernment matters more. Most of it is just misinformation, after all. People say knowledge shouldn't be restricted, but now we have the opposite problem. There's so much information that just skimming through it takes too much time. On top of that, as we shift from text to video, getting information has become even harder. Compared to text, YouTube videos feel like they have much lower information density. I've heard that the TikTok generation's text literacy is declining, but maybe that's actually a social adaptation to process as much data as possible from low-density sources In that sense, the efficiency of RAG ultimately comes down to what kind of good knowledge you're feeding into the AI.
- Philpax 3mo agoThe title suggested a far more interesting piece than the actual post. Alas.
- hx8 3mo agoI don't understand why Open Knowledge Format improves interoperability, which is the main claim for its value. These LLMs are obviously advanced enough to navigate other MD file organization schemas like Obisidan, or other text files like Emacs Org.
- golly_ned 3mo agoHas anyone figured out why anyone would bother adopting the google 'open knowledge format'? Normally I expect a set of tooling to be build on top of any open format. Value-adds and interoperability. Instead I just see a way to organize markdown files.
- nezhar 3mo agoStill looking for this, the OpenKB project look promising.
- mikewarot 3mo ago>Personal wikis always died for the same reason. Mine (WikidPad) died when I switched to Linux, and learned that breaking changes to WxPython rendered it worthless, as none of the dialog boxes functioned after that point. Sure, the source from 2012 was available, but my Wiki really wasn't. Eventually this forced me back to Windows... but it was too late, now I'm back on Windows, and still don't trust WikidPad. >Keeping them current was tedious, and humans hate tedium. But the tedium is the one thing language models are immune to. They will happily re-link, re-summarize, and reconcile contradictions across a hundred files without complaining. Yeah, and you're going to trust the LLM to reliably maintain this? Not a wise choice. I really wish we had a reliable way to annotate and interlink documents using hypertext. Unfortunately, HTML doesn't actually let you mark up (annotate) hypertext. We still, 81 years later, don't have a Memex! 8(
- skeledrew 3mo agoI've been using org[0] for the past 10 years to store knowledge, project write-ups, notes, etc and love it. As I get deeper into AI I'm keeping that, like recently I created an org-edit tool to manage especially issues in the project write-ups. 1 file with all my several hundred projects accumulated over the years, and the value has only grown although it's becoming harder for me to personally consume; I'll likely just create a couple commands that improve the browsing experience. But I continue to love and prefer that single large file (actually several: personal knowledge, projects, business and an inbox) to many individual files. And it's synced across my devices, where esp on mobile I can access it all via Orgzly[1]. [0] https://orgmode.org/ https://orgmode.org/ [1] https://orgzlyrevived.com/ https://orgzlyrevived.com/
- skybrian 3mo agoIf you're maintaining Markdown with a coding agent then you need linting tools to do things like checking for broken links. Without consistency checks, it will often make mistakes. The more internal consistency checks you can do, the better. This is good practice anyway, and a coding agent can help write the tools. Now that we have coding assistance, we can even be more ambitious. A common, language-independent test suite would be more useful than Markdown for generating an SDK and then verifying that it matches the spec. So I don't think plain Markdown is the best way.