Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
n1xis10t
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
n1xis10t
2mo ago
What does the icon look like?
2.
▲
by
n1xis10t
2mo ago
Such a weird thing. .aigc isn’t a real toplevel domain, and when I searched for it I found that apparently TikTok uses it as an abbreviation for “ai generated content”. That fits with the fake headlines. I was going to suggest checking 
3.
▲
by
n1xis10t
2mo ago
Wow, so ~43 million a day, ~15 billion a year? That’s sick. How much space do you have on the server, and how many pages do you intend to scale to?
4.
▲
by
n1xis10t
2mo ago
Yeah that’s pretty weird. I thought that maybe it was websites getting hacked and then redirecting to malicious stuff, but then I did your recommended search and it all seems very benign, and the videos seem to match the page title and urls
5.
▲
by
n1xis10t
3mo ago
What will happen if you don’t use AI? Will they fire you?
6.
▲
by
n1xis10t
3mo ago
Was the reason for your attempt to search in this way that you got the “internal server error” message (that I just got) when you tried to use the website’s search feature?
7.
▲
by
n1xis10t
3mo ago
It seems pretty cool, but I don’t think I’m the best person to ask. Most of social media I don’t use, and I’m not terribly interested if the results are from the apis directly. If it was an independent index of these sites, or if it took a
8.
▲
by
n1xis10t
3mo ago
How does it work? Is it like a meta search engine where your service sends the query to all the apis, or do you actually have an index of all this stuff?
9.
▲
by
n1xis10t
4mo ago
I haven’t tried that, thanks for the tip. It is far from perfect of course, because there are tons of great high quality sites that use .com. What we really need is a good search engine that is the scale of google, but those don’t happen al
10.
▲
by
n1xis10t
4mo ago
Thank you for saving me from the misery and regret.
11.
▲
by
n1xis10t
4mo ago
Gotcha, it’ll be interesting to see how it progresses.
12.
▲
by
n1xis10t
4mo ago
Oh, the rules are that paged out articles are only one page, so a longer article would have to go somewhere else like 2600. The collection of about 1.073 million pages in extracted text form that I have takes up about 4.8 GiB spread across
13.
▲
by
n1xis10t
4mo ago
You’re Canadian? That’s pretty hilarious, I am too. It must be something they put in the Timbits. For promotion, I’d recommend picking the most technically interesting part of your implementation, something that’s really clever, and then ma
14.
▲
by
n1xis10t
4mo ago
Very cool, I subscribed to the newsletter. I’ve experimented with retrieval and ranking across a sample of a million pages from the early days of the Common Crawl (around 2014) and I was surprised by how many of them seemed high quality. Th
15.
▲
by
n1xis10t
4mo ago
Would you mind writing a comment about this search engine you have found or created? I’m intrigued, but I don’t like clicking on links right away.
16.
▲
by
n1xis10t
4mo ago
Why?
17.
▲
by
n1xis10t
4mo ago
I just thought to try putting “music” in place of “www” in the playlist url, but unfortunately IA and CC still have nothing.
18.
▲
by
n1xis10t
4mo ago
Unfortunately I got the same result that archivarix did above, nothing in CC and nothing in IA. The thumbnail had a different link than the title of the playlist so I thought I’d try that, but the wayback machine redirects to the first vide
19.
▲
by
n1xis10t
4mo ago
I might not be able to help, but I’d like to give it a shot. Can I see the two IA links? If it was the “this page hasn’t been archived” error I’m less likely to be able to do anything, but I can check all the Common Crawl indexes and see if
20.
▲
by
n1xis10t
4mo ago
Update: So I mustered the courage to try the search engine, because it was looking not very much like a scam, and it becomes very apparent as soon as you use it that non-deleted videos are also indexed.
21.
▲
by
n1xis10t
4mo ago
Seems pretty cool. So this is a recent project, and you haven’t been working on this since 2005 right? Have you considered also indexing videos that haven’t been deleted?
22.
▲
by
n1xis10t
4mo ago
I like it. 10000 pages is very small, but it looks good and seems to function well. Is there any ranking? You might find this article interesting: https://archive.org/details/search-timeline It’s mostly about search en
23.
▲
by
n1xis10t
5mo ago
Does it support full text keyword search of the messages? Also do you have plans to scale to the size of searchcord (63 billion messages), or at least the recent discord message dataset (2 billion)?
24.
▲
by
n1xis10t
5mo ago
Sounds like blekko had a larger impact on the early urls than I thought. Out of curiosity, do you remember how large blekko’s index was at it’s peak?
25.
▲
by
n1xis10t
5mo ago
I think that the only other funding model (other than ads) that I’ve seen is a subscription, and I’ve only seen Kagi do that. It seems to work well for them though, last time I checked they had something like 50’000 subscribers. I kind of l
26.
▲
by
n1xis10t
6mo ago
It is 3 years old, so I highly doubt that it will work anymore. Google requires you to run JavaScript now, and there were some workarounds but they patched those. So unless it gets results from something else that uses Google, like Yahoo Ja
27.
▲
by
n1xis10t
6mo ago
That’s really funny. I think I’ve seen one of these repos before, but I don’t think I’ve used one and I haven’t created one. Not too late to hop on the bandwagon I suppose!
28.
▲
by
n1xis10t
6mo ago
Personally I just sit in a dark hole and wish that I knew things
29.
▲
by
n1xis10t
6mo ago
Are you sure you actually logged into tumblr and didn’t just put your login information into a phishing thing?
30.
▲
by
n1xis10t
6mo ago
So is the idea to use this to make a bunch of links to your site and boost its ranking in search engines? Does that work?
More ›