Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jhoetter
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
From failed proprietary product to funded open-source
(blog.kern.ai)
2 points
by
jhoetter
4y ago
|
1 comments
2.
▲
by
jhoetter
4y ago
I thought this could be interesting - sharing how we iterated from a no-code product for non-technical users to a developer platform, and from proprietary to open-source. For anyone who is currently facing similar challenges, this could be
3.
▲
Show HN: Standardizing NLP for a Modern ETL
(bricks.kern.ai)
6 points
by
jhoetter
4y ago
|
1 comments
4.
▲
by
jhoetter
4y ago
Awesome that you guys finally went open-source :)
5.
▲
by
jhoetter
4y ago
Very much needed. Spoke with many of the people listed on the repo, can confirm that they are not only open-source but generally very kind and trying to help. Using this for our current fundraise, and hopefully using this for our next one,
6.
▲
by
jhoetter
4y ago
Annotation platforms use Excel. I once received 25 files of separate Excel spreadsheets from a labeling service for 10k texts (short texts about product titles, e.g. "Sauvignon blanc" -> "wine"). Had to merge them, wh
7.
▲
by
jhoetter
4y ago
Nice, thanks. If you have any questions, please don't hesitate to contact us. Here's our Discord: https://discord.com/invite/qf4rGCEphW
8.
▲
by
jhoetter
4y ago
Thanks Francis! Means a lot :)
9.
▲
by
jhoetter
4y ago
My guess (some if this we already have, some we don't): - automation: integration of heuristics (multiple columns that you can program via formulas and such) - exploration: finding outliers or most similar records given some reference
10.
▲
by
jhoetter
4y ago
Thanks, means the world!
11.
▲
by
jhoetter
4y ago
I think so too. Mostly that it is something open. I also believe that it will change the workflow a bit, and that DVC will play a major role in it for versioning your different data hypotheses. Let's see, exciting times ahead!
12.
▲
by
jhoetter
4y ago
Hey Mark, I totally agree. We're focusing on NLP, but we're generally interested in what programming will develop into. To exaggerate a bit, but I like that idea: With "regular programming" (not the best term, but I mean
13.
▲
by
jhoetter
4y ago
Maybe also PowerPoint?
14.
▲
by
jhoetter
4y ago
Hi Tom! Thanks, happy to hear that :) We've focused on JSON as the user-specified data model. So you can upload anything fitting into a JSON. We're using pandas to process the uploaded data, so spreadsheets or CSV-ish also work. W
15.
▲
by
jhoetter
4y ago
Awesome, thanks for the suggestion. Already installed :)
16.
▲
by
jhoetter
4y ago
Hey, thanks! :) What exactly do you mean with application? You can just pull the repository. As mentioned in the installation section, it is quite easy to start the app on your local host.
17.
▲
by
jhoetter
4y ago
The most famous is arguably Snorkel, which started with an open-source library as a research project. We used that a lot ourselves, but the library by now is deprecated. We aim to extend on that idea by providing something that comes as clo
18.
▲
by
jhoetter
4y ago
Looks interesting, I'll check it out!
19.
▲
by
jhoetter
4y ago
Hi Ruben, you can take a look at our architecture overview here: https://github.com/code-kern-ai/refinery#-architecture A bit below it, you find a table with the links to all repositories. All of them are open-source.
20.
▲
by
jhoetter
4y ago
Hey, I'm Johannes - one of the maintainers of refinery. Thanks Jonathan for sharing!! Would be super excited if you guys have any feedback. It's nowhere near perfect yet, but you can already use it to build some great data-centric