36 ms·
I scraped all of OpenAI's Community Forum
- ritavdas 2y ago[flagged]
- olesya1979 2y ago[dead]
- xandrius 2y agoLove it, just for the sole reason of turning something OpenAI made into a dataset for everyone else :D
- codetrotter 2y agoI don’t think OpenAI are gonna lose any sleep over this. Isn’t a “community forum” like this basically just: “we’re not gonna spend money on providing adequate customer support so instead here is a forum where y’all can talk amongst yourselves and we’ll give you some badges and imaginary points for doing the customer support yourselves”?
- solardev 2y agoThey probably just sic a customer service GPT on it and use it to train the other ones...
- alt-glitch 2y agoI believe a community forum is absolutely vital for an "ecosystem" company. There needs to be a town square where people can discuss ideas and share feedback about that particular ecosystem. OpenAI has a pretty active forum with moderators replying and helping out all the time.
- deleted 2y ago[deleted]
- varelse 2y ago[dead]
- miduil 2y agoThat's an interesting write-up, I wonder how this would look for other big Discourse communities such as NixOS.
- alt-glitch 2y agoThis is definitely a workflow we can package into something open-source. I wonder how the community moderators would like it.
- dcreater 2y agoI for one would love it!
- velid0 2y agoNow train a gpt based on the data :D
- testfrequency 2y agoBut make sure to call it ClosedData or something so we know it’s not open source (sorry, I think openai and sam are gross)
- davely 2y agoMaybe I don’t understand this sentiment, but are people really that hung up on the name? I see this sort of thing posted a lot (i.e., “it should be ClosedAI instead of OpenAI, lol”) What if it just means “Open for Business” instead of “Open Access for All”? Or maybe they should just make it an acronym? I’m sorry for the confusion on my part, but there’s just been a lot of words dedicated toward expressing frustration with the company because they chose to use “open” in their name. Personally, I don’t find it frustrating that Apple doesn’t sell fruit and Intel doesn’t actually give intelligence data.
- rootusrootus 2y agoIs the frustration because of the name, or because open [access] was part of their ethos at the beginning, and people think they've abandoned it?
- startupsfail 2y agoOpenAI is supposed to be a nonprofit. But, when the nonprofit board tried to exercise control, it became very clear that the nonprofit arm is not, in fact in control any longer. The board was wiped out, nearly everyone in the company seemingly was willing to join Microsoft or Sam Altman or what not. This doesn’t seem to be compatible with continuing loftily call themselves with the same name, as the initial nonprofit mission.
- woopsn 2y agoIt's a gimmick. When the nonprofit was organized in 2015, the name certainly did not mean open for business. It meant (loftily) undertaking the quasi-religious quasi-humanist mission "in the spirit of liberty" to generate a new kind of super wealth as "broadly and evenly distributed as possible". As in prepare for the end... THE END OF HIGH PRICES! > to benefit humanity as a whole, unconstrained by a need to generate financial return - https://openai.com/blog/introducing-openai https://openai.com/blog/introducing-openai
- throwaway98797 2y agodid they have the right to use all thier data? /s
- SunlitCat 2y agoI didn't even knew they have community forums. Looking at the main homepage (openai.com), the only external links I can find are to chatgpt and their docs hosted on platform.openai.com. The other links lead to their socials, github and soundcloud (of all places). Maybe I'm not looking thoroughly enough, so I may be wrong, tho!
- hughesjj 2y agoI would also love to see these forums both to post and to lurk
- enonimal 2y ago> Number of Posts with negative sentiment, grouped by Topic > # 1 Result: Python Packaging Checks out
- minimaxir 2y agoA pro-tip for using the OpenAI API is to not use the official Python package for interfacing with it. The REST API documentation is good, and just using it in your HTTP client of choice like requests is roughly the same LOC without unexpected issues, along with more control.
- rockostrich 2y agoI've found this happens with a lot of first party clients. At work, we use LaunchDarkly for feature flags and use their code references tool to keep track of where flags are being referenced. The tool uses their first party Go client to interact with the API but the client doesn't handle rate limiting at all even though they have rate limiting headers clearly documented for their API.
- klooney 2y agoFirst party clients are typically an afterthought, and you can't add features without getting a PM to sign off, which strangles the impulse to polish & sand down rough edges.
- rattray 2y agoAgreed. Any in particular come to mind that you'd like to see improved? (my company provides first-party clients with a lot of polish; maybe we could help)
- rattray 2y agoHey minimaxir, I help maintain the official OpenAI Python package. Mind sharing what issues you've had with it? (Have you used it since November, when the 1.0 was released?) Keen for your feedback, either here or email: alex@stainlessapi.com
- wavyknife 2y ago(disclaimer: I work for Discourse) Discourse has an AI plugin that admins can run on their community to generate their own sentiment analysis (among other things), though it's not quite as thorough as this write up! https://meta.discourse.org/t/discourse-ai-plugin/259214 https://meta.discourse.org/t/discourse-ai-plugin/259214 We're always interested to see how public data can be used like this. It's something that can be a lot more difficult on closed platforms.
- Aachen 2y ago> helps you keep tabs on your community by analyzing posts and providing sentiment and emotional scores to give you an overall sense of your community for any period of time [...] > Toxicity can scan both new posts and chat messages and classify them on a toxicity score across a variety of labels Is that within the defined data processing purposes of all Discourse setups? Does the tool warn admins they might need to update their policies before being able to run this tool, perhaps needing to seek consent (depending on their jurisdiction and ethics)? It sounds somewhat objectionable, trying to guess my mental state from what I write without opt-in Edit: and apparently it also tries to flag NSFW chat messages, does Discourse have PM chats where this would flag private messages for admins to read or is it only public chats that this bot runs on? > tagging NSFW image content in posts and chat messages
- BadHumans 2y agoMore companies and communities than you think already do this without your knowledge let alone consent.
- david_allison 2y agoThat doesn't mean we can't do better
- BadHumans 2y agoBetter at what though? I don't even think it's a problem to begin with.
- dorkwood 2y agoI did a bit of data scraping for fun in the past, but I was never quite sure of the legality of what I was doing. What if I was breaking some law in some jurisdiction of some country? Was someone going to track me down and punish me? OpenAI has taught me that no one gives a shit. Scrape the entire internet if you want, and use the data for whatever you feel like.
- ifyoubuildit 2y agoDo you think it would be better if someone did track you down and punish you? Which world do you want to live in?
- n0sleep 2y agoI think large companies should be punished for stealing from people to make themselves richer.
- EcommerceFlow 2y agoA precursor to this would have been that Linkedin lawsuit Microsoft lost, allowing that one company to scrape all of Linkedin (technically "public information").
- htrp 2y agohiQ Labs v. LinkedIn
- alt-glitch 2y agoWe were really heading someplace with The Semantic Web aka The Real Web 3.0 [1] Alas we have to fight against the machines in order to properly read the internet thru machines. I believe Discourse knowingly keeps its data easy to scrape though, so kudos to them! [1]: https://en.wikipedia.org/wiki/Semantic_Web https://en.wikipedia.org/wiki/Semantic_Web
- bsuvc 2y ago> OpenAI has taught me that no one gives a shit. Scrape the entire internet if you want, and use the data for whatever you feel like. Cloudflare gives a shit. My household had to use our 5G internet for most things for a week or two until our IP reputation recovered.
- xfalcox 2y agoThat's super cool, thanks for sharing! I will share this as an easy to follow example of what we can with AI. > Allowing a Q&A interface using these embeddings over the post contents could speed up research over the community posts (if you know the right questions to ask :P). Let's view some posts similar to this one complaining about function calling That's indeed a great thing to surface, and that's exactly how the the OpenAI forum selects the "Related Topics" to show at the end of every topic. We use embeddings for this feature, and the entire thing is open-source: https://github.com/discourse/discourse-ai/blob/main/lib/embeddings/semantic_related.rb#L13 https://github.com/discourse/discourse-ai/blob/main/lib/embe... We also embeddings for suggesting tags, categories, HyDE search and more. It's by far my favorite tech of this new AI/ML gen so far in terms of applicability. > Using Twitter-roBERTa-base for sentiment analysis, we generated a post_sentiment label (negative, positive, neutral) and post_sentiment_score confidence score for each post. We do the same, with even the same model, and conveniently show that information on the admin interface of the forum. Again all open source: https://github.com/discourse/discourse-ai/tree/main/lib/sentiment https://github.com/discourse/discourse-ai/tree/main/lib/sent... Disclaimer: I'm the tech lead on the AI parts of Discourse, the open source software that powers OpenAI's community forum.
- deleted 2y ago[deleted]
- fzysingularity 2y agoSo epic, thank you for making this dataset available to everyone!
- warlord1 2y ago[dead]
- klooney 2y agoWhat's the "Day Knowledge Direction" cluster in the Atlas view?
- alt-glitch 2y agoNeat find! That's actually a cluster of all the system messages notifying users about closing and re-opening of the thread. That's why they're so tightly clustered. I believe the naming isn't perfect for this, but this was all automatic topic modelling! Example: [1]: https://community.openai.com/t/read-this-before-posting-a-new-question/51 https://community.openai.com/t/read-this-before-posting-a-ne...
- garyiskidding 2y agoThis is really amazing. Pretty insightful. Thank you.
- alright2565 2y agoI saw this part: > Every Discourse Discussion returns data in JSON if you append .json to the URL. then this: > Raw data was gathered into a single JSONL file by automating a browser using Playwright. Kinda seems to me like having a whole browser instance for this isn't necessary? I would have been surprised if this .json pattern didn't continue for all pages, and it turns out that it does in fact also work for the topic list: https://community.openai.com/latest.json https://community.openai.com/latest.json The other place I've seen this sort of API pattern is reddit. For example, https://www.reddit.com/r/all.json https://www.reddit.com/r/all.json or (randomly chosen) https://www.reddit.com/r/mildlyinfuriating/comments/1bqn3c0/after_19_years_my_ipod_nano_seems_to_have_kicked.json https://www.reddit.com/r/mildlyinfuriating/comments/1bqn3c0/...